Skip to content
Blue Headline

Search Blue Headline

Or browse a topic

Technology · AI & Robotics

AI Browser Agents Claimed 90% Success. Real Websites Told a Different Story

Operator and Mariner are gone, and agents now live in Chrome, Claude and ChatGPT. An independent test shows how often they still fail.

Blue Headline ·

AI Browser Agents Claimed 90% Success. Real Websites Told a Different Story

Blue Headline explains research plainly. The sources are linked below, with how much to trust them.

Updated September 2026: rewritten to reflect big changes since March, when two of the four agents we compared have since shut down.

AI browser agents promise something simple. You describe a web chore, and the AI clicks, types and fills in forms for you.

In March, we compared four of them. Six months later, the picture looks very different.

OpenAI’s Operator and Google’s Project Mariner have both shut down. Their features now live inside ChatGPT, Gemini and Chrome.

So here’s an updated guide: what’s still around, how well these agents really work, and the risk that matters most.

  • Browser agents are now built into tools many people already use, including Chrome, Claude and ChatGPT.
  • An independent test found most web agents completed only about 30% of real tasks.
  • A hidden trick called prompt injection can turn an agent against its own user.

What is an AI browser agent?

An AI that uses the web like you do

A browser agent is an AI that can see a web page and act on it. It clicks buttons, scrolls, types into fields and moves between tabs.

That’s different from a chatbot that only tells you what to do. An agent reaches for the mouse itself.

Why people want one

Much of our online time goes on small, dull steps. Filling in the same forms, comparing prices across tabs and copying details from one site to another.

A good agent could take those steps off your plate. A bad one could make expensive mistakes with total confidence.

What changed since March

Here’s how the four agents from our original comparison have moved on.

Agent in MarchWhat happenedWhere the feature lives now
OpenAI OperatorShut down on 31 August 2025Agent features inside ChatGPT
OpenAI’s Atlas browserLaunched October 2025, being retired from 9 August 2026Browser features moving into ChatGPT and Codex
Google Project MarinerShut down on 4 May 2026Gemini’s agent features and “auto browse” in Chrome
Anthropic computer useStill a tool for developersPlus Claude in Chrome, generally available since 26 August 2026
Perplexity CometFree to download since October 2025Now on Windows, Mac, Android and iOS

The pattern is clear: standalone agent experiments are fading. The features are moving into browsers and assistants people already have open.

The main options today

Gemini in Chrome with auto browse

Google added “auto browse” to Chrome in January 2026. It can handle multi-step chores such as filling in forms, scheduling appointments and shopping within a budget.

It’s available to Google AI Pro and Ultra subscribers in the US, on Windows, Mac and Chromebook Plus.

Google says it’s designed to pause and ask for your confirmation before sensitive actions, such as making a purchase or posting on social media.

Claude in Chrome

Anthropic‘s Chrome extension became available on every paid Claude plan on 26 August 2026. It reads pages, clicks links, types, fills in forms and works across tabs, using your existing logins.

It now approves actions it judges safe automatically, checked by a safety classifier. You can switch that off and approve every step yourself.

Perplexity Comet

Comet is a full browser built around Perplexity’s AI assistant. It became free to download in October 2025 and now runs on Windows, Mac, Android and iOS.

Its assistant can browse, fill in forms and complete multi-step tasks. Some advanced features depend on a paid Perplexity plan.

ChatGPT’s agent features

OpenAI folded Operator into ChatGPT as an agent that can use a browser. It is now also retiring its separate Atlas browser and moving those features into ChatGPT itself.

OptionWhere it runsWho can use itBuilt-in safeguard
Chrome auto browseInside ChromeUS Google AI Pro and Ultra subscribersPauses before purchases and posts
Claude in ChromeChrome extensionAny paid Claude planSafety classifier; manual approval optional
Perplexity CometIts own browserFree to downloadAgent features vary by plan
ChatGPT agentInside ChatGPTPaid ChatGPT plansAsks before important actions

How well do they really work?

The independent test

Company demos always look smooth. So researchers at Ohio State University and UC Berkeley built a tougher test called Online-Mind2Web.

It has 300 realistic tasks on 136 live websites. The study was presented at the COLM 2025 conference.

A reality check

Many agents had claimed success rates of around 90% on an older benchmark. On the new test, most managed about 30%.

Agent (as tested in 2025)Tasks completed
OpenAI Operator61%
Claude computer use (3.7)56%
SeeAct (research agent)31%
Browser Use30%
Claude computer use (3.5)29%
Agent-E28%

Only the two newest agents cleared 50%. The authors warned that earlier results painted an overly optimistic picture.

Browser agents now live inside Chrome, Claude and ChatGPT. But independent testing found even the best completed only about 6 in 10 real web tasks.

Why this matters now

Those tests are from 2025, and today’s agents are newer. They are very likely better.

But nobody has published an equally careful independent test of the current versions. Until someone does, treat any claim of near-perfect reliability with caution.

Where browser agents still break

Even good agents fail in predictable ways.

FailureWhat it looks likeWhat to do
Layout changesClicks the wrong button after a site redesignWatch high-stakes steps
Losing the threadForgets the goal halfway through a long taskSplit big jobs into smaller ones
OversteppingActs where it should only lookKeep approval on for payments and deletions
False confidenceSays “done” when it isn’tCheck the result yourself

Websites also change constantly. An agent that handled a site perfectly last week can stumble after a redesign.

That’s one reason tests on live websites matter more than polished lab benchmarks.

My rule of thumb: use agents to remove friction, not to replace your judgement.

The biggest risk: prompt injection

Hidden instructions on web pages

Prompt injection is when a web page hides instructions for the AI. You ask the agent to summarise a page, and hidden text tells it to do something else.

Because the agent acts with your logins, that “something else” could reach your email, bank or work accounts.

It has already happened

In 2025, security researchers at Brave showed that Perplexity’s Comet could be tricked by text hidden on a page. The hidden text told it to fetch one-time passcodes from the user’s email.

Another team later reported a Comet flaw it called “CometJacking”. Perplexity says it has since fixed the issue.

The defences are improving

Anthropic says that, in its own testing, attacks against an older Claude model with safeguards succeeded 16.7% of the time. For its newest models, it reports close to zero.

Those are the company’s own numbers, not an independent audit. But they show how seriously the industry now takes the problem.

If you want to go deeper, our guides on trusting AI agents and MCP server security cover the same risks from other angles.

What earlier research found

Benchmarks can flatter agents

The Online-Mind2Web study’s main lesson was about testing itself. Older benchmarks used easier, more static tasks, which made agents look far more capable than they were on the live web.

It also found that automated grading can be unreliable. The team built an AI judge that agreed with human reviewers about 85% of the time.

Security is a category-wide problem

Brave’s follow-up research found similar prompt-injection weaknesses in other AI browsers, not just Comet. It described the issue as a systemic challenge for the whole category.

How it fits together

Agents are getting more capable and more widely available. But reliability and security are improving more slowly than the marketing suggests.

How much should you trust this?

Early. Product facts come from the companies, and independent tests lag behind new releases.

What makes it convincing

  • Every product change here comes from official announcements or help pages.
  • The success rates come from a peer-reviewed conference study using live websites.
  • The security examples were documented by independent researchers.

What makes me cautious

  • The independent success rates are from 2025 versions of these agents.
  • Safety numbers for the newest models come from the companies themselves.
  • Products in this area change every few months.
  • Features and availability vary by country and subscription.
This guide showsThis guide does not show
Which agents exist today and where they liveWhich agent is best for your exact tasks
How agents performed in an independent 2025 testHow today’s versions perform in the same test
That prompt injection is a real, documented riskThat any agent is fully safe from it
Practical ways to reduce riskGuarantees about future products

What this means for you

Browser agents are worth trying for low-stakes chores. Just don’t hand them the keys to everything.

  • Start with boring, low-risk tasks. Research, comparing options and filling in simple forms are good first jobs.
  • Keep approval on for money and deletions. Anything costly or irreversible should need your click.
  • Use a separate browser profile. Keep the agent away from your banking and main email where you can.
  • Watch the first few runs. You’ll quickly learn where your agent slips.
  • Be suspicious of “summarise this page” on unknown sites. That’s exactly where hidden instructions lurk.

In this IBM Technology video, an expert explains how prompt injection attacks work:

What we still don’t know

  1. How reliable are today’s agents? A fresh independent test of current versions is overdue.
  2. Can prompt injection be solved? Defences are improving, but no one has declared victory.
  3. Who’s responsible when an agent errs? The rules for mistaken purchases or deletions are unclear.
  4. Will standalone AI browsers survive? Two big players have already folded theirs into existing products.
  5. How will websites respond? Some may block agents, others may build pages designed for them.

My take: useful helpers, not trusted assistants yet

What I like about this moment is that agents have become practical. You no longer need a special experiment to try one.

I’m less convinced by claims that they can run your digital life. The best independent evidence still shows plenty of failures, and prompt injection is a genuine threat.

For now, I’d treat a browser agent like a capable new intern. Give it clear, low-risk jobs, and check its work before anything important goes out the door.

Main source: An Illusion of Progress? Assessing the Current State of Web Agents

Last reviewed: 2026-09-28

Covers: AI browser agents from Google, Anthropic, Perplexity and OpenAI, their reliability and prompt-injection risk

Evidence: Early — product facts from official announcements; independent reliability data from 2025 versions only

Share this

Comments

No comments yet. What did you think?

The Blue Headline briefing

The best stories. Only when they matter.

Free, no account needed, unsubscribe anytime. We only send when it’s actually worth reading.