Home › Agents & Tool Use › Computer / browser use
🕹️ · Build

Computer / browser use

Driving a GUI or browser by screenshots and clicks when no API exists.

In one line

The agent looks at a screen, decides where to click or type, and repeats until the task is done, automating software that offers no programmatic interface.

ConceptWhat it is

Computer use, also called browser use, is an agent pattern where the model operates ordinary software the way a person does, reading the screen and issuing mouse and keyboard actions instead of calling a structured API. The agent receives a rendered screenshot or an accessibility / DOM tree, reasons about what it sees, and emits low-level actions like click at these coordinates, type this text, scroll, or press Enter. A controller executes those actions against a real browser or desktop, the screen changes, and the loop repeats.

It exists to reach the large surface of systems that have no usable API: legacy enterprise desktop apps, internal web tools behind logins, vendor portals, and anything a human can do in a browser but a developer cannot script cleanly. Because the interface is the pixels and widgets themselves, the same agent can in principle work across any application, at the cost of the brittleness and slowness that come from acting through a UI rather than a contract.

How it worksThe mechanics

A goal is given in natural language; the agent captures the current state as a screenshot and/or structured accessibility tree, the model plans a single next action, and a driver, a browser-automation library or OS input layer, executes it as a click, keystroke, scroll, or navigation. The environment updates, the agent re-observes the new state, and it repeats this observe, plan, act loop step by step, often re-grounding on each frame because element positions can move. The loop ends when the model judges the goal met, a stopping condition or step budget is hit, or a verification check confirms the outcome.

At a glanceSee it

Computer / browser use diagram
Computer / browser use diagram 1

The grounding problem — how a high-level intent like click the refund button becomes an actual on-screen coordinate, through raw pixels, a UI tree node, or a numbered set-of-marks overlay.

Computer / browser use diagram 2

A safety gate the bare loop omits — screen text is screened for injected instructions, and consequential actions like purchases are held for human confirmation before they commit.

When to use itWhere it fits

  • The target system has no API, SDK, or stable integration point and exposes only a human-facing UI.
  • You must automate across many heterogeneous web apps or legacy desktop tools without building a connector for each.
  • Tasks are exploratory or long-tail, filling forms, navigating portals, or gathering data, where bespoke scripting is not worth it.
  • A human can supervise and the workflow tolerates occasional retries and slower runs.

When NOT to use itLimits & anti-patterns

  • A clean API or direct data access exists, call it instead; it is faster, cheaper, and easier to verify.
  • The task is high-volume or latency-sensitive, where per-step screenshots and model calls are too slow and costly.
  • Errors are unacceptable or irreversible, such as payments or production changes, without strong guardrails and confirmation.
  • The UI changes constantly or has strong anti-automation defenses like CAPTCHAs and bot detection.

Trade-offsAdvantages & costs

Advantages
  • Works anywhere a person can click, needing no integration, vendor cooperation, or API.
  • One general capability generalizes across arbitrary apps instead of one connector per system.
  • Unlocks automation of legacy and closed systems that would otherwise require manual labor.
  • Mirrors human workflows, so runs can be demoed, recorded, and reasoned about the way a person would.
Trade-offs & costs
  • Brittle: small layout, timing, or wording changes break the run, and pop-ups or modals derail it.
  • Slow and expensive: every step is a screenshot plus a model call, so multi-step tasks accumulate many high-cost turns.
  • Hard to verify: success is judged from pixels, so silent failures and wrong-but-plausible actions slip through.
  • Safety exposure: an agent with live keyboard and mouse control can take real, hard-to-undo actions.

ExampleIn the real world

An operations team needs to reconcile invoices from a supplier portal that offers no export API. A computer-use agent is given the goal "download last month's paid invoices and enter their totals into the finance sheet." It opens the portal in a controlled browser, logs in with supplied credentials, reads the screenshot to locate the date filter, clicks and types the range, waits for the table to render, opens each invoice, and reads the total from the detail view. For each one it switches to the spreadsheet tab and types the value into the next row. A human reviews a short summary before the sheet is saved, and any invoice the agent could not read confidently is flagged rather than guessed.

ToolsHow to implement it

  • Anthropic's Claude computer use, a capability where the model takes a screenshot and returns mouse and keyboard actions to drive a virtual desktop or browser.
  • OpenAI's Operator and its underlying Computer-Using Agent (CUA), which drive web pages by reading the rendered view and acting on it.
  • browser-use, an open-source library that pairs an LLM with Playwright so agents can navigate and act on web pages.
  • Playwright and Selenium, the underlying browser-automation drivers that actually execute the clicks, typing, and navigation.

Cost & effortWhat it takes

This is one of the most expensive agent patterns per task. Each step sends a fresh screenshot, a large image input, plus accumulated context, and non-trivial tasks run dozens to hundreds of steps, so token and call counts climb quickly and latency is measured in minutes rather than seconds. Vision-capable models and a sandboxed browser or VM add infrastructure cost. Engineering effort shifts from building integrations to building reliability: retries, step budgets, verification checks, human review, and monitoring. Budget for high variable inference cost and ongoing maintenance as target UIs drift, and reserve the pattern for cases where the absence of an API rules out cheaper automation.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Agentic features are reaching mobile surfaces, so an agent's integration target can be a phone app rather than a desktop or API.

    TechCrunch AI · 23 Sep 2026 · source

  • Updated this page Platform terms, not just technical capability, can block an agent from completing a purchase on a user's behalf, so agent checkout flows need per-site permission rather than assumed open access.

    The Verge AI · 20 Sep 2026 · source

  • Updated this page OpenAI has shipped GPT-6 Astra, a new flagship business model with computer use and stronger reasoning, which the model boards do not yet list.

    Add GPT-6 Astra to the model fact file and boards as a new frontier model with computer use, and update the newest-entry as-of date.

    OpenAI · 9 Sep 2026 · source

  • Updated this page Claude Cowork now runs in Anthropic's Chrome extension, extending agentic workflows to in-browser tasks with skills and plugins.

    Anthropic · 12 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning