How AI Agents Browse the Web: From Fetch to Click

An AI agent can reason, write code, and call tools. But the web is where most real-world information lives. So how does an agent actually use a web page? The answer has evolved through three distinct generations, and which one your agent uses determines what it can and can’t do.

Generation 1: Fetch and Parse

The simplest approach. The agent makes an HTTP request, converts the HTML to text or markdown, and reads it. No browser, no JavaScript, no clicking. This is what most coding agents ship today. OpenCode has two tools in this category:

  • webfetch — a direct fetch() call that converts HTML to markdown via turndown. It spoofs a Chrome User-Agent to avoid bot detection and falls back to an honest UA if Cloudflare blocks it. A 5MB response cap and 30-second timeout keep it bounded.
  • websearch — calls Exa’s MCP API (https://mcp.exa.ai/mcp) via JSON-RPC. This is a gateway: Exa does the actual crawling and indexing, the agent just receives text results.

Claude Code does the same with its WebFetch tool, with one addition: it can pass the fetched content to a model for inline summarization. You give it a URL plus a prompt, and it returns the model’s answer about the page rather than the raw page itself. Useful for large pages that would blow the context window. Cline and Roo Code ship similar fetch-and-convert tools. Cursor takes a different angle — web access is built into the editor UI, not exposed as an agent tool. The user triggers it via @web mentions, and it routes through Cursor’s backend. Note this generate can not do javascript rendering, only read static html.

Generation 2: Accessibility Tree Navigation

The agent drives a real browser (Chromium, Firefox) but doesn’t look at pixels. Instead, it reads the page’s accessibility tree as structured text and interacts with elements via stable references. This is what Hermes Agent (by Nous Research) does, and it’s the sweet spot for cost versus capability.

The agent calls a sequence of tools, reads text, decides what to click. Mostly it’s realized by parse the DOM tree, but also can use vision based LLM processing as a back-up, it sure consumes a lot of tokens. Hermes has a drive_preview tool that returns only a delta after each action — what was added, removed, or changed — instead of the full page again. Safety would be a huge concern: a logged-in browser session is dangerous. It can make purchases, send messages, change account settings. Hermes treats browser sessions as “credential-bearing state” and gates destructive actions — purchases, messages, account changes, uploads — behind explicit confirmation.

Generation 3: Vision and Coordinates

The most capable and most expensive approach. The agent takes screenshots, the LLM processes them as images, and decides where to click by coordinates.

Devin is the primary example. It has a full browser automation tool with screenshots, clicks, and typing. It needs this because it’s designed for autonomous web tasks on arbitrary UIs it has never seen before — not just coding.

Most agents live at fetch-and-parse. Hermes lives at the accessibility tree — the sweet spot. Devin goes all the way to vision because it must handle arbitrary UIs autonomously. If you were to implement it yourself, the sample simple codes:

# browser_agent.py — simplified Hermes-style browser tool
from playwright.async_api import async_playwright
class BrowserTools:
"""Accessibility-tree-based browser automation for LLM agents."""
def __init__(self):
self.browser = None
self.page = None
self._ref_map: dict[str, str] = {} # @e1 -> CSS selector
async def navigate(self, url: str):
if not self.browser:
pw = await async_playwright().start()
self.browser = await pw.chromium.launch(headless=True)
self.page = await self.browser.new_page()
await self.page.goto(url)
await self.page.wait_for_load_state("networkidle")
return await self.snapshot()
async def snapshot(self) -> str:
"""Get the accessibility tree as text with element refs."""
tree = await self.page.accessibility.snapshot()
self._ref_map.clear()
lines = []
self._walk_tree(tree, lines, prefix="")
return "\n".join(lines)
def _walk_tree(self, node, lines, prefix, depth=0):
if not node:
return
role = node.get("role", "")
name = node.get("name", "")
if node.get("focused") or role in ("button", "link", "textbox", "checkbox", "menuitem"):
ref = f"@e{len(self._ref_map) + 1}"
# Store a way to find this element later
self._ref_map[ref] = name # simplified; real impl uses more robust selectors
lines.append(f"{prefix}[{ref}] {role} \"{name}\"")
for child in node.get("children", []):
self._walk_tree(child, lines, prefix + " ", depth + 1)
async def click(self, ref: str):
selector = self._ref_map.get(ref)
if not selector:
return f"Unknown ref: {ref}"
await self.page.get_by_role("button", name=selector).first.click() # simplified
return await self.snapshot()
async def type(self, ref: str, text: str):
selector = self._ref_map.get(ref)
if not selector:
return f"Unknown ref: {ref}"
await self.page.get_by_role("textbox", name=selector).first.fill(text)
return f"Typed '{text}' into {ref}"
async def screenshot(self) -> bytes:
return await self.page.screenshot()

The engines underneath

None of this works without a real browser doing the rendering. The two libraries that make it possible are Puppeteer (Google, 2017) and Playwright (Microsoft, 2020) — built by the same team, Playwright being the successor. Both talk to a real browser via the Chrome DevTools Protocol, but Playwright won the agent ecosystem for three reasons: it has a first-class Python client (agents are increasingly Python, not Node), it drives Chromium/Firefox/WebKit with one API (useful for evading bot detection), and it auto-waits for elements to be actionable before clicking (agents don’t know when a page is “ready”). Hermes, Devin, Aider, Browser Use, and Browserbase all run on Playwright under the hood. The libraries aren’t agent tools themselves — they’re the layer between the agent’s tool calls and the real browser process, translating browser_click("@e3") into page.click(selector) and returning the result as text the LLM can read.

Leave a Reply