TL;DR

Training web agents with RL requires browser infrastructure that can run reliably at scale. Models will continue to become better at navigating the web. The bottleneck then shifts from capability to infrastructure and identity. At Browserbase, we have built the framework for secure and reliable web authorization and accessibility and are excited to partner with HUD to run frontier RL environments and evals that labs need.

What a browser eval is

An unpredictable, dynamic web is core to what makes browser evals a valuable performance metric for models. A reproducible browser eval becomes valuable in discerning agent capability, requiring the LLM to standardize success across changing websites.

What a task is

A task encompasses a prompt, an environment with the files and tools the agent can use, and an outcome grader. QA agents additionally review execution traces for prompt alignment and reward hacking, flagging instances where an agent is given or withheld credit. This gives teams a way to identify misleading scores and support reliable, verifiable grader outputs.

Beyond outcomes, The LiveWeb Samples we used in this blog are an example of a task set which groups them together for an eval run. Each sample uses a browser environment, sandbox and configuration (name, starting URL, prompt, rubric).

See the HUD overview, Tasks & Tasksets, and Creating Environments.

Common pitfalls in browser evals

Sufficient evidence and reproducibility are imperative in distinguishing model errors from shifts in websites, infrastructure, or grader-related discrepancies.

Given that the web itself is a live environment, it serves as an uncontrollable variable between runs. As a result, the opportunity for agents to reward hack exacerbates across several form factors. Improvements in browser and computer use enable models to modify web pages and deceive graders that check the appearance of success versus underlying outcomes. Agents can also access publicly available task set solutions, thus increasing success rates without verifying any sincere performance improvements. In this same vein, Anthropic has documented cases where models have exhibited eval awareness, recognizing that it was running BrowseComp and retrieving the benchmark’s answer key instead of solving the original research problem.

Identity and authorization also bear significant weight against an agent’s success on the web. Bot detection and CAPTCHAs often gate agents from accessing the intended task environment. Verified mode leverages Web Bot Auth to sign traffic and enables access providers like Cloudflare and reCAPTCHA to recognize agents. This allows website owners to set policy and enable broader access for agents.

LiveWeb: frontier-grade tasks on the open web

HUD’s LiveWeb tasks contain browser tasks that are hard for frontier models, and they have open-sourced three tasks to give us a look into how they design frontier level problems. These tasks cover camera control, reading photos and driving a data tool.

TaskWhat's
Examine a 3D ladybird model embedded on Sketchfab and count the spots on its shell, head excludedone integer; exact count scores 1, closer misses score higher
On diamondrosesanctuary.com, a vacation rental site, find how many chairs are at the kitchen tablethe exact count
Use EPA's AirNow interactive map to find air quality data near zip 35173 for April 19, 2024site / site ID / pollutant / daily AQI (separate weights)
Sign in to an existing Costco account using the supplied credentials and confirm the session is authenticatedsign-in / correct credential entry / confirmed login / accurate outcome report without unrelated account actions (25% each)

All four are read-only with answers that sit on stable websites. Ladybird grading takes one final integer with the exact count scoring a 1 and closer misses scoring higher than farther ones. Ranges or a list of alternatives do not count. The task asking about the number of chairs at the kitchen table requires an exact-match, because noticing the partially hidden chair is the visual challenge. AirNow grades four values with separate weights, so finding the right site but missing the AQI value still produces a score.

Taskset: https://hud.ai/tasksets/1fc0f549-5830-4061-bc3a-25ad502ce723.

Task definition

LiveWeb samples runs these tasks on one shared HUD environment. The agent starts on a live URL and works from a prompt. Each task is YAML data under tasks/ - name, starting URL, prompt, and a rubric used only for grading, not a new environment.

(https://github.com/hud-evals/liveweb-samples):

@env.template(id="web-research")
async def web_research(prompt: str, url: str, rubric_items: list[dict]):
    await load_browser_on_url(url)
    answer = yield prompt
    yield await grade_with_rubric(answer, rubric_items)

The environment opens the starting URL, gives the agent the prompt, then scores the answer against the rubric. To add a task, add those fields under tasks/ and sync. You do not rebuild the browser environment unless the task needs a new archive.

Where models fail

We ran all four tasks across model tiers.

ModelLadybirdChairsAirNowCostco Login
GPT 6 Astra93.3%0.0%100.0%75%
Claude Fable 5.156.0%0.0%100.0%75%
Claude Sonnet 50.0%0.0%100.0%75%
Qwen 3.8 Max0.%16.7%95.0%75.0%

On the Ladybird task, Astra’s six attempts average 93.3%. Fable averages 56% over five completed attempts. Sonnet and Qwen sit at 0 on completed Ladybird runs; Qwen’s Ladybird comparison is also confounded, because that harness had shell and files but no computer tool, and several runs asked for a URL or image instead of counting.

AirNow is mostly solved when the run finishes. Twenty of twenty-one completed AirNow attempts score 1.0 as recorded. Astra’s 50% cell includes three provider capacity errors that recorded 0; on its three completed AirNow runs it scores 100%. Qwen’s 95% includes one 0.7 that looks like a grader parsing bug (judge text says MET, parser stored UNMET). Keep 95% until that is fixed.

Chairs is almost all zeros under an exact-match grader that expects 7. One Qwen run scores 1.0 after installing object detectors.

Authenticated websites are an example where Browserbase helps production agents get past access issues. In this task, the agent needs to sign into an existing account by entering credentials without exposing the password in tool output, and validate that it's able to access the authenticated account page. HUD uses Browserbase’s cloud browser sessions, live-site navigation, screenshots and a secure way to enter credentials in these cases to validate agent performance on real websites.

Frontier labs target these failure points when they collect data. A task that causes models to fail at a specific step, is more valuable than one they simply pass or fail overall. This makes the full trajectory the most valuable component.

Monetizing evals on Datavendor

Frontier labs need data to improve their models, so they buy private evals, environments, and tasksets that current models find difficult. A taskset that meets the standard above, with graded traces as proof, becomes a monetizable asset.

"Historically, data has been something where you're in the room. You're in San Francisco, you have these long-term relationships. But it's not just engineers in San Francisco who can make frontier-grade training data. There are experts around the world who are not in the room." - Parth Patel, co-founder of HUD, on the Standard Capital podcast

Use hud deploy to publish the environment and hud sync to publish the tasks, then list them on datavendor.ai.

What’s next

To learn the principles behind good task design, read What makes a good task.

Take the liveweb-samples repo, sync it into your taskset on HUD, and run it directly from the platform. Or eval from the CLI against a public taskset:

hud eval liveweb-samples claude --runtime hud --all --max-steps 150 -y

For the stack around this blog:

Start building with Browserbase

Run headless browsers for your agents and automations at scale. Get started free in minutes.

Sign up for free