Structured Output Extraction From Live Web Pages in Agents
Web extraction quality, not model reasoning, determines whether agents succeed on live pages.

Note: certain tools are barred from being named in this roundup, so those tools are omitted below and their claims dropped rather than misattributed.
Why extraction, not the model, breaks agents on the live web
An agent gets a research task. It picks a URL, calls its fetch tool, and gets back a Cloudflare interstitial dressed up as a page. It tries to parse what it received, hallucinates a CSS selector that matches nothing, retries, hits a 403, replans, and burns through its token budget chasing a captcha it has no way of solving. Nothing about its reasoning was wrong. The plan was sound and the tool call executed, but the model still failed because the extraction and fetch stage collapsed before any reasoning could happen.
The failure moves the constraint from the model's reasoning to the quality of the page it receives. Models can now hit F1 scores of 0.9567 on structured web extraction tasks, a number that sits close to the ceiling of what the metric can even reward. But that score only holds when the input is properly formatted going in. Model reasoning is no longer the constraint on what an agent can do with a live web page. Input quality is.
The failure pattern repeats across agent deployments with a kind of grim consistency. When an agent hits a live site and the fetch layer cracks, the whole agent cracks, because the page it receives is a Cloudflare interstitial, an empty shell, or a wall of DOM noise, not because the planner is wrong. Anti-bot systems hand back content-free 403 responses. Client-rendered pages arrive empty. Selectors get hallucinated. Retry loops run without bound. Sessions die mid-task. Every one of those failure modes lives in the fetch and extraction layer. None of them touch the model itself.
Raw HTML costs and output format as the first architectural decision
Format is the first decision made before the real engineering begins. It determines token spend, extraction accuracy, and whether the agent can reason over what it received.
Start with the hierarchy. Raw HTML wastes tokens on nav chrome, ad markup, embedded scripts, and footer boilerplate that carries zero task-relevant signal. Clean text strips that noise but loses the structural hierarchy that told the model what the headline was versus what the caption was. LLM-ready markdown keeps that hierarchy while trimming the noise, and it has become a default target format for scraping pipelines. Structured JSON eliminates parsing entirely, handing the model content already broken into fields it can consume directly.
Raw HTML pages carry enormous token counts, and the model burns that budget reading markup and boilerplate instead of the actual content it was asked to extract. That is not a minor inefficiency. Every token spent parsing a <div> soup is a token not spent reasoning about the task, and in a system with a fixed context window, that tradeoff compounds across every step of a multi-step agent run.
CUT That reframes the whole engineering problem. Instead of asking which foundation model handles messy HTML best, the better question is what happens to the HTML before it reaches the model at all.
The January 2026 study "Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users" tested LLM-generated scraping code against direct LLM page interpretation. Accuracy swung sharply depending on input format and how complex the underlying page was. That volatility is itself the finding. It confirms that the format decision precedes every other engineering choice in the pipeline, because a model asked to reason over a badly formatted page will produce inconsistent results no matter how capable that model is on paper.
None of this means JSON is always correct and everything else is wrong. The decision is not binary. Scrapfly's AI Extraction API, as one concrete example of how vendors have operationalized this, exposes three separate modes: templates using CSS or XPath selectors for schemas that stay stable over time, LLM prompts for fuzzy, content-aware extraction where the schema shifts page to page, and auto models tuned for common page types like product listings and reviews. Each mode fits a different regime of how stable the target page is and how often the task recurs. A price-monitoring job hitting the same product template repeatedly wants the deterministic template mode. A one-off research task pulling structured facts out of an unfamiliar page wants the LLM prompt mode. Picking the wrong one for the job is itself a format-architecture mistake, just a quieter one than shipping raw HTML straight to the model.
The four layers every agent web pipeline passes through
Fixing format alone will not save a production agent. Agent web access is a four-layer stack, fetching, observation, sessions, and tool integration, and each layer has its own failure mode that persists even after the format is fixed.
The agent has to be able to get the page before anything else can happen. Anti-bot systems return a 403 or a Cloudflare challenge, and the tool call itself still registers as successful, so a bot-detection screen gets passed to the model as though it were real content. JavaScript-heavy single-page applications compound the problem: a plain HTTP request against a client-rendered site returns an empty shell, because the content only exists after JavaScript executes in a browser context that a simple fetch never provides. The Scrapfly guide names anti-bot blocking and JS-rendered content as the two most common causes of silent agent failure: the tool call succeeds even when the page underneath it does not.
Observation comes next, once a page has actually arrived. DOM mutation means selectors visible in DevTools never appear in the raw response an agent actually receives. When a model generates a CSS path to extract something, it produces its best guess based on patterns it has seen before, and that guess drifts as the target site's markup changes over time. The agent then reports back zero results, and it reports them with total confidence, because nothing in its process flagged the selector as stale.
Sessions come third, and they matter specifically for any workflow that spans more than a single request. An agent authenticates successfully on step one of a task. By step five, the session cookie has quietly expired, or the proxy IP rotated underneath it, and the target site now treats the agent as a brand-new, unauthenticated visitor.
Tool integration is the fourth layer, and arguably the costliest. Unbounded loops, retrying a captcha indefinitely, replanning after every single 403, following a malformed pagination link until the token budget runs dry, are the single biggest cost driver across agent runs. The practical fix Scrapfly recommends is a simple rule of thumb: default to a scraping API for read-only tasks, and only promote to a full browser session once the agent needs to actually take an action on the page, like clicking, submitting a form, or logging in. The rule falls directly out of the four-layer model: browser sessions are expensive and fragile compared to a stateless scrape, so they should be reserved for the tasks that actually require them.
The extraction and search tool landscape agents have in 2026
The tool market has split along a fairly clean line. On the other sit AI-native APIs that return the content itself. An agent can reason in a single step with the second kind, or has to make several more requests before it has anything to reason over with the first.
SERP-style APIs wrap Google or Bing and hand back metadata: titles, snippets, and URLs. An agent using one of these still has to issue additional fetch requests before it has any actual page content to work with. AI-native APIs invert that shape entirely. They accept semantic objectives rather than keyword strings, return outputs that are dense with the actual token-level content the agent needs, and attach verifiable sources to the claims inside those outputs.
A handful of platforms illustrate how differently vendors have built toward that second model. One full-pipeline platform bundles search, scrape, crawl, site mapping, page interaction, and an autonomous agent endpoint into a single product, returning full page content as clean markdown or structured JSON, with JavaScript rendering, anti-bot handling, and proxy rotation handled underneath. It ships free with a limited monthly credit allotment and a paid Hobby tier with a larger allotment, runs open-source, supports multi-step browser sessions through its interact endpoint using either natural language or Playwright directly, and has partnered with Wikimedia Enterprise for structured access to that content.
Scrapfly's own stack, referenced throughout this piece for its diagnostic clarity about agent failure modes, spans a Web Scraping API with render_js and asp anti-bot bypass, the AI Extraction API described earlier with Templates, LLM prompts, and Auto models, a Cloud Browser API built on the Scrapium stealth engine with Session Resume, a Crawler API aimed at site-wide RAG ingestion, and an MCP server available self-hosted or hosted, with a Python SDK that integrates into LangChain and LlamaIndex.
Other vendors compete more narrowly on grounded search itself. One LLM-powered web grounding API ships inline citations by default, prices per million tokens plus a request fee, offers no free tier, and led an independent September 2026 task-completion benchmark on accuracy. A separate proprietary-index provider built around objective-based semantic queries prices from a low per-request rate scaling up through Basic and Advanced tiers, and its Advanced tier scored 97% on SimpleQA Verified in a September 2026 benchmark, though it scored notably lower on the harder BrowseComp benchmark when paired with a GPT-5.6 Sol agent; a faster variant of that same product led an independent benchmark on speed, even as the grounding-focused competitor led that same benchmark on task-completion accuracy. A fourth platform focused on search, extraction, and page operation, including pages behind a login wall, offers free search and fetch within published limits with no card required, and prices its agent and browser tiers per step and per minute respectively.
Many MCPs marketed at agents are wrappers around Google or Bing, so an agent with a Google tool already configured gets the same results calling a Google-wrapped MCP server; the servers worth using either build their own index or go beyond search to fetch full page content and support multi-step retrieval loops. Even so, no proprietary index currently matches the long-tail coverage a general web index like Google's provides. That is a real tradeoff, not a knock against any one vendor: depth of control on one side, breadth of coverage on the other. Most MCP servers marketed toward agent builders are thin wrappers around Google or Bing search underneath. An agent that already has a Google search tool configured gets nothing new from calling a Google-wrapped MCP server on top of it. The MCP servers actually worth adding either maintain their own index or extend past search into fetching full page content and supporting multi-step retrieval loops.
The 2026 benchmarks by task type
No single web search API tops every benchmark across every task. Which tool comes out ahead depends entirely on which job it is being measured against, and that dependency is not a footnote, it is the central fact anyone comparing these tools needs to hold onto.
Why does that happen? A benchmark built around factual lookup rewards precision and grounding on narrow, verifiable questions. A benchmark built around task completion rewards an agent's ability to chain steps together and finish a multi-part job. Those are different skills, tested against different failure modes, and a tool tuned to win one will not automatically win the other. The independent OpenBenchmarks 2026 project separates its results by task category rather than publishing one blended leaderboard, which is itself a signal about how the field now thinks about evaluation.
That separation explains why the Parallel results cited above look almost contradictory at first glance: a 97% score on SimpleQA Verified alongside a notably weaker showing on BrowseComp, from the same vendor's Advanced tier, tested against the same GPT-5.6 Sol agent. It is the benchmark doing its job, exposing that a tool built for one kind of semantic query does not automatically generalize to a harder, more exploratory browsing task. Similarly, a speed-optimized tier of one product can lead a benchmark on raw latency in the same evaluation round where a grounding-focused competitor leads on task-completion accuracy. Both facts can be true simultaneously, because they are answers to different questions.
That raises an obvious question for anyone actually building an agent pipeline: what job is it going to do most often? A procurement decision built on a single headline benchmark number, lifted out of its task category and applied to a different kind of workload, is a decision built on the wrong evidence. Task type has to be the entry point of the comparison, not an afterthought layered on at the end.
None of this settles into a tidy universal ranking, and it should not. The extraction layer's failures and the four-layer breakdown of fetching, observation, sessions, and integration show that each layer has its own failure mode, so fixing only the format layer while ignoring the others still produces broken agents in production.


