On a Thursday in late August 2026, Anthropic introduced Browser Use, a tool that grants its Claude model a more principled way of reading and acting upon the web — not by guessing at pixels on a screen, but by consulting the same structured map of a page that assistive technologies have long relied upon. The shift is modest in appearance yet significant in kind: it moves AI web interaction from visual approximation toward something closer to genuine comprehension of structure. In doing so, it raises the enduring question of whether better tools make us more capable, or simply more dependent on
Anthropic's Browser Use tool replaces pixel-guessing with structured page references
References point to the actual element, not where it appears on screen.
Why does it matter that Claude can use references instead of coordinates?
Because coordinates are fragile. A button at x: 640, y: 320 on one screen might be at a completely different position on another, or after the page reflows. References point to the actual element in the page structure, so they work regardless of where the element appears visually.
But doesn't the reference break too, if the page changes?
It does, but the system catches it. When Claude tries to use a stale reference, the executor rejects it and asks Claude to read the page again. It's a controlled failure, not a silent miss.
What's the real win here—is it just accuracy?
Accuracy, yes, but also efficiency. The bigger win is batching. Instead of Claude clicking a button, waiting for feedback, then filling a form, then submitting—three separate model calls—it can now request all three actions at once. That cuts latency and token costs dramatically on long workflows.
Why can't it just do all three at once anyway?
Because each action depends on the result of the previous one. If the click fails, the form field might not exist yet. The executor has to run them in order and watch for failures, but it only needs one round trip to the model instead of three.
Who actually runs the browser?
The developer does. Anthropic doesn't execute the browser actions. The developer's application translates Claude's requests into actual browser commands, maintains the session, and sends results back. That's different from the Skills API, where Anthropic runs the code in its own sandbox.
Does that create security problems?
It creates responsibility. Claude can encounter prompt injection in web content or be tricked into navigating somewhere dangerous. Anthropic recommends isolation—containers, virtual machines, minimal permissions. JavaScript should stay off unless you need it.
El Pulso
- Web automation powered by screenshot coordinates has always been fragile — layouts shift, elements move, and the AI is left guessing at a moving target.
- Browser Use replaces that guesswork with accessibility tree references, giving Claude the same structural map that screen readers use, making interactions more stable and predictable.
- Action batching collapses what were once many costly round-trips into a single model turn, cutting latency and token expenses as workflows scale into dozens or hundreds of steps.
- Security remains an open wound — prompt injection, unauthorized redirects, and the risks of JavaScript execution mean developers must isolate browser sessions in hardened environments.
- The tool enters a field already occupied by Playwright and Puppeteer, and while the concepts converge, the protocols diverge — leaving developers to weigh Anthropic's approach against adapting tools they may already trust.
On a Thursday in late August 2026, Anthropic introduced Browser Use, a tool that grants its Claude model a more principled way of reading and acting upon the web — not by guessing at pixels on a screen, but by consulting the same structured map of a page that assistive technologies have long relied upon. The shift is modest in appearance yet significant in kind: it moves AI web interaction from visual approximation toward something closer to genuine comprehension of structure. In doing so, it raises the enduring question of whether better tools make us more capable, or simply more dependent on the architecture others have built.
Anthropic has released Browser Use, a tool that changes how Claude navigates the web. Rather than analyzing screenshots and estimating pixel coordinates — a method prone to drift and error — Claude now receives a structured representation of the page drawn from its accessibility tree, the same underlying map that screen readers use. Each interactive element carries a labeled reference, and Claude acts on those references directly, bypassing the visual guesswork that made earlier approaches unreliable.
The release arrived alongside Computer Use, the Skills API, and the Files API moving into general availability. Browser Use is accessible through the Claude API and brings with it a default set of 27 browser operations. When Claude calls read_page, it receives a text rendering of the page's structure with each element tagged. If the page changes and a reference goes stale, the executor must catch the mismatch and prompt Claude to re-read before continuing — the API will not handle this automatically.
Efficiency was also a target. Previously, a simple sequence — click, fill, submit — required separate model calls with feedback between each step, consuming tokens and adding delay. Browser Use allows Claude to batch multiple actions into a single turn. The executor carries them out in order and returns all results together, reducing round-trips and lowering costs as workflows grow more complex.
The implementation carries meaningful constraints. Browser sessions live in the developer's own environment, not Anthropic's infrastructure, placing the burden of session management and security squarely on the developer. Anthropic recommends isolated containers with JavaScript execution and file uploads disabled by default, since code Claude generates runs with the page's own privileges. Batched actions also complicate approval workflows — a routine early action in a sequence might eventually lead somewhere requiring human sign-off, so executors must monitor and pause as needed.
The tool enters a landscape with established alternatives. Playwright already exposes accessibility snapshots and element references through its MCP server, and Puppeteer offers comparable building blocks. An unrelated open-source project shares the Browser Use name. The conceptual overlap is real, but the protocols differ, and developers must decide whether Anthropic's specific approach justifies building around its own client-toolset architecture rather than extending what already exists.
Anthropic has released a new Browser Use tool that fundamentally changes how Claude interacts with web pages. Instead of analyzing screenshots and guessing at pixel coordinates—a method prone to error and inefficiency—Claude now receives a structured map of the page itself, complete with labeled references to buttons, links, text fields, and other interactive elements. When Claude wants to click something, it no longer has to calculate where that button sits on the screen. It simply references the element by its assigned label, like ref_3, and the browser executes the action directly.
The tool arrived Thursday as part of a larger Anthropic release that also brought Computer Use, the Skills API, and Files API into general availability. Developers can access Browser Use through the Claude API using the browser_toolset_20260801 identifier. The underlying mechanism draws on the accessibility tree—the structured representation of a page that screen readers and assistive technologies use to navigate websites. By giving Claude access to this same tree, Anthropic has created a more reliable and efficient way for the AI to understand and manipulate web content.
This shift from coordinates to references addresses a real problem in web automation. Computer Use, Anthropic's broader desktop automation tool, works by looking at screenshots and sending mouse and keyboard commands tied to specific pixel positions. That approach works across an entire desktop but struggles with web pages, where layouts shift, elements move, and visual appearance alone cannot reliably tell Claude what something is or how to interact with it. Browser Use operates within the browser itself, where the underlying structure is available and stable. When Claude calls read_page, the executor returns a text representation of the accessibility tree, with each interactive element tagged with a reference. If that reference becomes stale—if the page navigates or changes significantly—the API will not automatically catch the mismatch. The executor must recognize when a reference no longer points to a valid element, reject the action, and ask Claude to read the page again before proceeding.
Anthropnic has also tackled the efficiency problem that emerges when AI agents interact with browsers. Previously, if Claude wanted to perform a sequence of actions—click a button, fill in a form field, submit the form—it would have to wait for feedback after each step before requesting the next one. That meant multiple round trips to the model, each one consuming tokens and adding latency. The new version allows Claude to request multiple browser actions in a single turn. The executor carries them out in order, since each action depends on the state created by the previous one, and returns all the results together. This batching cuts unnecessary model calls, lowering both latency and token costs, particularly as workflows grow from a handful of interactions to dozens or hundreds.
The implementation details reveal important constraints. Browser Use is currently available only through the Claude API and not yet inside Claude Managed Agents. The default toolset exposes 27 browser operations and adds roughly 6,600 input tokens to a request before screenshots and accessibility trees are even counted. Developers can disable operations they do not need to reduce that overhead. Unlike the Skills API, which runs code inside Anthropic's sandbox, or the Files API, which stores documents on Anthropic's servers, browser sessions remain in the developer's own environment. The developer's application must translate each of Claude's requests into an action inside its browser, maintain the session between turns, and return enough information for Claude to understand what happened.
Security considerations loom large. Claude can still encounter prompt injection attacks embedded in web content or be redirected to unexpected locations. Anthropic recommends running the browser in an isolated container or virtual machine with minimal access. JavaScript execution and file uploads should remain disabled unless explicitly needed, since code Claude generates runs with the page's privileges and can access data or make requests available to that page. Batching complicates approval workflows slightly, since multiple actions can arrive together and a routine click early in a sequence might eventually lead somewhere that requires user permission. The executor must check actions as they execute and pause for approval when necessary.
The tool arrives in a landscape where similar capabilities already exist. Playwright, Microsoft's browser automation library, can represent pages as ARIA snapshots and locate elements by role rather than coordinates. Microsoft's Playwright MCP server already exposes structured accessibility snapshots with references. The concepts align closely with Browser Use, but the protocols do not—Playwright speaks MCP while Anthropic uses its own client-toolset protocol, so developers would need an adapter to translate between them. Puppeteer offers similar building blocks through its Accessibility.snapshot() API and browser control functions. An open-source project also called Browser Use runs AI agents against Chromium through Chrome DevTools Protocol, but it has no connection to Anthropic's tool and would require separate integration work. For developers, the question becomes whether Anthropic's approach offers enough advantage to justify building on its specific protocol rather than adapting existing tools.
Citas Notables
Anthropic recommends running the browser in an isolated container or virtual machine with minimal access, since Claude can encounter prompt injection attacks and unauthorized redirects.— Anthropic documentation