An AI agent that opens a browser, clicks through a website and finishes the job for you is no longer a demo. In 2026 you can turn one on inside Chrome, call one through an API, or run one in a cloud workspace. The honest picture, though, is more mixed than the launch videos: these agents are good at short, well-defined web tasks, still struggle with long ones, and carry a security risk that no vendor claims to have fully solved. This guide covers what works today, what the research shows, and a practical setup that keeps the risk small, with a tested policy gate in Python you can adapt.
What a computer-use agent actually does
The loop is simple. The model receives a screenshot (or a structured view of the page), decides on one action (click here, type this, scroll, press a key), your software performs it, takes a new screenshot, and sends it back. Repeat until the task is done or the agent gives up. Everything the agent knows about the page comes from what it sees, which is exactly why a web page can influence it.
Where the products stand (October 2026)
- Claude in Chrome (Anthropic) became generally available on all paid Claude plans on August 26, 2026, after a year of gradual rollout. It can now act on its own instead of asking before every step, with a safety classifier checking each action, and enterprise admins can restrict it to approved domains.
- Anthropic's computer use API lets developers build their own agents. The current toolset no longer needs a beta header, and it runs on the Claude API, Google Cloud, Amazon Bedrock and, in beta, Microsoft Foundry.
- OpenAI now offers browser and computer use inside ChatGPT Work, announced in July 2026, which replaced the earlier ChatGPT agent. Its standalone Atlas browser, launched in October 2025, has since been deprecated according to OpenAI's own page.
- Google offers a computer use capability in the Gemini API (still labeled preview), and ended its Project Mariner experiment in 2026, moving the technology into Gemini and Chrome's "auto browse" feature, which reached Android in the US for AI Pro and Ultra subscribers in August 2026.
- Perplexity's Comet browser has been free for everyone since October 2025.
Availability changes quickly and often depends on plan and country, so check the provider's current page before you build on any of these.
How good are they? What the benchmarks say
OSWorld (2024) was the first widely used test: 369 real computer tasks across apps and websites. When it was published, humans completed 72.4% of them and the best model just 12.2%. Two years later, vendors report scores above 80% on its cleaned-up version, so the original test is close to solved, at least as reported by the model makers themselves.
That's why the same research group released OSWorld 2.0 in June 2026: 108 long, realistic workflows that take a person about 1.6 hours each and need around 318 actions, against roughly 30 in the first version. The results are sobering. The best model tested completed only 20.6% of tasks fully (54.8% with partial credit), and none of the agents completed any of the longest tasks. The authors describe the failures in plain terms: agents lose track of constraints, miss information that arrives mid-task, guess instead of asking the user, and skip checking their own work.
Vendors have since published higher OSWorld 2.0 numbers for newer models, but on partial-credit scoring, so they aren't comparable with the paper's full-completion rate. The practical reading: short, clearly defined tasks work well now; long, multi-hour workflows mostly don't, yet.
The real risk: prompt injection
A browser agent reads everything on the page, including text you can't see. If a page contains instructions written for the agent, it may follow them. This isn't theoretical:
- Brave's security team (August 2025) showed that a hidden comment on a Reddit page, plus a simple "summarize this page" request in Perplexity's Comet, could make the agent fetch the user's email address and a one-time login code from their Gmail and send them out. In October 2025 they showed the same class of attack using text hidden inside images, affecting several AI browsers, and concluded that agentic browsing remains inherently risky until the underlying design changes.
- Anthropic's own red-teaming of Claude in Chrome found a 23.6% attack success rate without defenses and 11.2% with its first mitigations (August 2025). Newer models and layered defenses have cut this sharply. On the test set Anthropic used for the August 2026 launch, attacks against the older Claude Opus 4.5 still succeeded 16.7% of the time with safeguards on, versus 0% to 0.3% for its 2026 models.
- OpenAI wrote in December 2025 that prompt injection, "much like scams and social engineering on the web, is unlikely to ever be fully 'solved.'"
The improvement is real, but low on a vendor's test set isn't zero on the open web, and attackers adapt. The safe assumption is that any page the agent reads might try to steer it, and the setup should limit what a steered agent can do. (Our guide to AI agent security and prompt injection explains the underlying "rule of two".)
A safe setup, from the vendors' own guidance
Anthropic's and Google's developer documentation give consistent advice. Combined with the consumer products' built-in safeguards, it adds up to six rules:
- Isolate it. Run developer agents in a dedicated virtual machine or container with minimal permissions. For browser extensions, use a separate browser profile with no banking, health or government accounts signed in.
- Allowlist the sites. Limit the agent to the domains the task needs. A link to anywhere else is blocked, not followed.
- Keep secrets out. Don't let the agent type passwords, card numbers or one-time codes. When a login is needed, the person takes over for that step.
- Confirm consequential actions. Purchases, sending messages, submitting forms, deleting things and accepting terms need a human click. Both Anthropic and Google build confirmation steps into their tools for this reason.
- Never solve CAPTCHAs. Google's documentation tells agents never to attempt to solve or bypass them. A CAPTCHA is a signal to hand control back to the person.
- Watch, log and limit. Keep a step budget, record every action, and watch the first runs of any new task. Anthropic's docs note that these agents are better suited to background tasks than real-time, human-paced work, so plan for that.
A tested demo: a policy gate in code
The safest place for these rules is in code that sits between the model and the browser, because text on a page can argue with a prompt but not with an if statement. The script below needs no browser or API key. It checks each action an agent proposes against an allowlist, a list of actions that always need confirmation, a list the agent may never do, and a step budget. The plan is scripted: the user asked the agent to re-order printer paper, and step 4 is what a hidden instruction on the supplier page tried to trigger.
"""A policy gate for a browser agent (no API key or browser needed).
Every action the model proposes passes through check() before it runs.
The rules live in code, so text on a web page can't talk its way past them.
"""
from urllib.parse import urlparse
ALLOWED_DOMAINS = {"supplier-portal.example", "docs.example"}
ALWAYS_CONFIRM = {"purchase", "submit_form", "send_message", "delete"}
NEVER = {"enter_password", "enter_card", "solve_captcha", "download_run"}
MAX_STEPS = 30
def domain(url):
host = urlparse(url).hostname or ""
return host[4:] if host.startswith("www.") else host
def check(action, step):
kind, url = action["kind"], action.get("url", "")
if step > MAX_STEPS:
return "STOP", f"step budget of {MAX_STEPS} used up; report progress to the user"
if kind in NEVER:
return "BLOCK", f"'{kind}' is never done by the agent; hand control to the user"
if url and domain(url) not in ALLOWED_DOMAINS:
return "BLOCK", f"{domain(url)} is not on the allowlist"
if kind in ALWAYS_CONFIRM:
return "ASK", f"needs the user's confirmation: {action.get('summary', kind)}"
return "ALLOW", "read-only or low-risk"
# Scripted plan: the user asked the agent to re-order printer paper.
# Step 4 is what a hidden instruction on the supplier page tried to trigger.
plan = [
{"kind": "open", "url": "https://supplier-portal.example/orders"},
{"kind": "read", "url": "https://supplier-portal.example/orders/1182"},
{"kind": "add_to_cart", "url": "https://supplier-portal.example/item/a4-paper"},
{"kind": "open", "url": "https://free-gift-claim.example/login"},
{"kind": "enter_password","url": "https://supplier-portal.example/login"},
{"kind": "solve_captcha", "url": "https://supplier-portal.example/checkout"},
{"kind": "purchase", "url": "https://supplier-portal.example/checkout",
"summary": "pay 240 AED for 10 boxes of A4 paper"},
]
for i, action in enumerate(plan, 1):
verdict, why = check(action, i)
print(f"{i}. {action['kind']:14} {verdict:5} {why}")
print("\nstep 31:", *check({"kind": "read", "url": "https://docs.example/a"}, 31))
Output, run on 2026-10-07:
1. open ALLOW read-only or low-risk
2. read ALLOW read-only or low-risk
3. add_to_cart ALLOW read-only or low-risk
4. open BLOCK free-gift-claim.example is not on the allowlist
5. enter_password BLOCK 'enter_password' is never done by the agent; hand control to the user
6. solve_captcha BLOCK 'solve_captcha' is never done by the agent; hand control to the user
7. purchase ASK needs the user's confirmation: pay 240 AED for 10 boxes of A4 paper
step 31: STOP step budget of 30 used up; report progress to the user
What the gate did:
- Normal browsing went through: opening the order history, reading an old order, adding paper to the cart.
- The injected detour was blocked: the "free gift" site isn't on the allowlist, so the agent never reaches it, however convincing the instruction was.
- Passwords and CAPTCHAs went back to the person. These are hand-off moments, not agent work.
- The payment paused for confirmation, with a plain-language summary of what would be bought and for how much.
- The step budget stops runaway loops and makes the agent report progress instead of wandering.
Real products implement richer versions of the same ideas (site permissions, action classifiers, confirmation prompts). If you build your own agent, write this layer first.
What these agents are good for today
- Good fits: gathering information from several sites into one table, filling repetitive web forms from data you provide (with a final review), checking order or shipment status, routine admin in web dashboards, testing your own website's flows.
- Poor fits for now: anything involving money without a human check, tasks that need your logins on sensitive accounts, long multi-hour workflows, and sites with heavy anti-bot protection.
Costs matter too. Every step sends a screenshot (Anthropic estimates roughly 1,000 to 1,800 tokens each, plus the tool definition), so a 50-step task adds up. Prefer an API or a structured connector over screen-clicking whenever one exists; it's faster, cheaper and harder to manipulate. (More on keeping agent bills down in how to cut AI agent costs.)
The takeaway
Browser and computer-use agents crossed from novelty to useful in 2026, for short, well-scoped tasks. They aren't yet reliable for long workflows, and prompt injection is a risk you manage, not one you eliminate. Give the agent a separate profile or machine, a short list of allowed sites, no secrets, and a human click before anything that spends money or can't be undone. Then start with one small, boring, repetitive task and watch it work.
Sources: Xie et al., "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024); "OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks" (June 2026); Anthropic, Claude for Chrome announcement (August 2025) and Claude in Chrome general availability (August 26, 2026); Anthropic computer use tool documentation; Google Gemini API computer use documentation; Brave security research on indirect prompt injection in AI browsers (August and October 2025); OpenAI, post on hardening ChatGPT Atlas against prompt injection (December 2025) and ChatGPT Work announcement (July 2026). Checked October 2026; products change quickly.