Ilustración editorial para BrowseComp: qué mide un agente que encuentra un dato difícil en la web y por qué su acierto no prueba que haga investigación fiable
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

BrowseComp targets a specific capability: retrieving a difficult fact

BrowseComp is a benchmark for evaluating agents that navigate the internet in search of hard-to-locate factual information. Its design starts from a simple observation: a question whose final answer fits in a few words may require a long chain of searches, reformulated queries, page openings, and connections among dispersed facts. The difficulty therefore does not necessarily come from writing a lengthy explanation or solving a mathematical problem; it comes from finding suitable evidence on a heterogeneous web.

According to the documentation and the introductory paper, the dataset contains 1,266 problems. Each seeks a short answer that can be checked against a reference. This choice reduces a common ambiguity in research evaluations: automatically judging a long response requires deciding whether its arguments, sources, and qualifications are sufficient. In BrowseComp, the outcome is closer to checking whether the agent reached the requested fact.

That purpose matters for teams comparing search agents. A system that solves this type of task has shown, at least under that protocol, an ability to sustain goal-directed exploration and retrieve a specific fact from intertwined information. However, that evidence does not automatically establish that it can perform high-quality open-ended research in the editorial, analytical, or business sense of the term.

The distinction matters because, in common usage, “research” usually includes more operations than locating an answer. It may involve framing a still-ambiguous question, identifying relevant sources, explaining conflicts among them, assessing date and authority, citing traceably, summarizing uncertainty, and deciding when evidence is insufficient. BrowseComp does not aim to cover all of those operations through its short-answer metric.

02

The evaluation unit simplifies grading, not search

The basic structure of a task separates two things that should not be confused. On one side is the process: the agent searches, browses, and decides what information to retain. On the other is the evaluated outcome: a short final answer compared with a reference answer. The fact that the outcome is brief does not mean the search trajectory is trivial; the benchmark was designed precisely so that relevant information is hard to find and requires persistent browsing.

The verifiability of the final answer is a methodological advantage. It makes it possible to calculate an accuracy rate without asking a human evaluator to read thousands of reports. It also limits the metric’s scope. If an agent gets a bare answer right, the result does not by itself reveal which pages it consulted, whether it interpreted its evidence correctly, or whether it could explain its reasoning to a user.

Nor should a match with the reference always be treated as equivalent to well-grounded research. An agent may reach the answer through prior knowledge, an incidental clue, or a robust search; the final score can be identical. Later work such as LiveBrowseComp raises precisely the need to distinguish evidence-based search from merely verifying something the system already appears to know. That question does not invalidate BrowseComp, but it narrows how a correct answer should be interpreted.

Conversely, a failure does not necessarily prove a lack of research capability. The web may change, a link may stop working, a search engine may alter its index, or a page may become subject to an access restriction. In an open-web evaluation, the result combines the agent’s ability with the state of its infrastructure and external resources at the time of the run.

03

A published percentage belongs to a system and protocol, not only to a model

It is tempting to summarize a result as though it were a stable property of a model. For browsing agents, that simplification often hides decisive variables. The result comes from a composite system: a model, search and page-reading tools, instructions, working memory, an exploration policy, time or action limits, and a mechanism for producing the final answer.

The available search engine can change which documents are retrieved and in what order. A browser with limited rendering may not access the same content as a full browser. Request limits, domain restrictions, cookie handling, or location can alter the viable path. Token budget, maximum number of steps, and the stopping rule also change results: an agent that can continue searching for longer has more opportunities to recover a decisive clue, but it may also become unfocused.

Another less visible variable is how many trajectories are allowed per question. One figure may result from a single run; another from several independent attempts with voting, selection, or later aggregation. These setups answer different questions. The first is closer to performance in a single interaction. The latter may measure the performance of a pool of samples and a selection strategy. Neither is inherently wrong, but they are not interchangeable.

For that reason, before comparing model cards, vendor announcements, or an internal team’s results, it is worth requesting the complete harness. Responsible comparison begins by establishing whether the conditions that generated both percentages are materially equivalent.

Minimum matrix before comparing two BrowseComp figures

VariableWhat should be documentedWhy it changes the interpretation
Version and item setEdition used, any exclusions, and execution datePrevents distinct datasets or runs from being treated as identical
Web accessLive web, cache, snapshot, or closed corpusDetermines what evidence was available
ToolsSearch engine, browser, extraction, limits, and domainsChanges page retrieval and reading
BudgetTime, steps, tokens, queries, and requestsAffects the practical depth of exploration
SamplingOne trajectory, multiple attempts, voting, or selectorChanges the statistical meaning of the percentage
GradingOutput format, normalization, and error treatmentDefines what counts as a correct answer
04

The live web makes the benchmark relevant, but complicates reproducibility

BrowseComp measures navigation on the internet, not merely querying a frozen database. That decision has a clear advantage: it retains some of the friction faced by a real agent. Answers may require reaching poorly visible pages, relating mentions, or persisting through unhelpful results. A fixed corpus would remove part of that dynamic and could make the task resemble conventional document retrieval more than web browsing.

The cost is that the web is not a stable environment. Pages are updated or disappear; search indexes are reordered; blocks, rate limits, and paywalls arise; answers may vary by region, language, or personalization. Even with no change to the model, a new run may not have access to the same clues as an earlier run. A historical figure should therefore be read alongside the date, tools, and evaluation incidents.

This tension cannot be resolved merely by declaring one mode superior to the other. An evaluation on the live web retains ecological validity for current browsing, but reduces repeatability. An evaluation with a frozen corpus or snapshot facilitates auditing and comparison, but leaves out real changes in availability and discovery. Both can be useful if they precisely describe what they measure and what they sacrifice.

BrowseComp-Plus is presented as a different proposal, aimed at a more transparent and controlled evaluation of deep-research agents. It should not automatically be treated as a new measurement on the same scale, nor should its results be added to or compared with BrowseComp results without examining the tasks, sources, protocol, and grading rule. A shared name is not a substitute for methodological equivalence.

LiveBrowseComp likewise raises a complementary issue: when questions concern recent facts, evaluation can help test whether an agent is searching for available evidence or merely reproducing prior knowledge. Its materials describe a set of 335 questions and mechanisms intended to reduce dataset leakage. That approach can provide another signal, but it remains a distinct benchmark rather than an automatic update to BrowseComp results.

Protocol for preserving the traceability of a run

  1. 01Fix the benchmark version, date, and list of items actually evaluated.
  2. 02Record the model, system instructions, tools, search engine, access limits, and regional configuration.
  3. 03Define the budget, number of trajectories, stopping policy, and aggregation rule before running the evaluation.
  4. 04Save final answers, error status, tool traces where permitted, and the reason for excluding each item.
  5. 05Separate agent failures, infrastructure failures, and unevaluable items in the report.
  6. 06Repeat a sample when the live web is part of the protocol and report the variation observed.
05

What can be inferred from a high score

A high score, obtained under well-documented conditions, is evidence that the evaluated system was able to retrieve a large number of short, difficult answers correctly within that dataset. In particular, it is reasonable to treat it as a signal of search persistence, an ability to turn a question into exploration, and skill at connecting clues to reach a specific factual finding.

It can also be an operationally relevant signal for products whose work ends with that kind of retrieval. For example, an internal workflow that needs to locate a specific fact and then submits it to human validation may benefit from an agent that finds better leads with less intervention. The useful test, however, is one that reproduces the sources, constraints, and consequences of the workflow itself, not only an external figure.

These inferences should be stated conditionally. They concern the system, tools, and budget used in the run. They do not justify attributing the result exclusively to the underlying model, nor turning it into a precise prediction of performance on an unknown distribution of real queries. The official documentation already warns that the short-answer format does not represent an open distribution of user queries.

In particular, a final-answer benchmark should not be turned into evidence about route quality. If a product requires auditability, its acceptance criteria should require the agent to return the sources it consulted, the evidence connecting them to its conclusion, and explicit treatment of limitations. Those properties may correlate with correctness, but they are not demonstrated by it.

06

What remains outside the metric, and why it matters in production

A BrowseComp score does not sufficiently measure whether an agent selects primary sources when they are available, distinguishes a competent source from an unreliable copy, or presents citations that let a user verify the conclusion. Nor does it require a long, coherent explanation. An agent may get an isolated fact right and still produce a flawed synthesis when it has to integrate multiple claims, dates, or definitions.

Ambiguity is another central limitation. Many real queries do not have a single answer without context: “largest,” “current,” “official,” “cost,” or “best” require specifying scope, date, jurisdiction, unit, or criterion. In a short-answer benchmark, ambiguity is reduced by constructing an evaluable reference. In production, a good agent must detect missing information, ask, or present alternatives, rather than optimize only for one final text string.

Currency and disagreement among sources require dedicated testing. A system may locate a difficult historical fact and fail when faced with information that changed yesterday. Likewise, it may retrieve a published claim without assessing that another source contradicts it. The neutrality of a synthesis, coverage of relevant perspectives, and handling of documentary conflicts are dimensions separate from finding an exact answer.

Finally, BrowseComp does not establish safety in subsequent actions. Browsing to seek information is not equivalent to being authorized or prepared to submit forms, make purchases, modify records, handle sensitive data, or execute business decisions. Those capabilities require specific controls, human validation proportionate to risk, and tests in the environment where they will be deployed.

Complementary tests by use-case risk

Product needTest that BrowseComp does not replacePractical criterion
Report with sourcesEvaluation of traceability and documentary relevanceEvery important claim should be linkable to accessible evidence
Ambiguous queryDataset containing incomplete or polysemous questionsThe agent asks for context or states alternative interpretations
Changing informationDated tests on recent informationThe agent communicates the verification date and detects staleness
Disagreeing sourcesCases with documented conflictThe agent represents disagreement without concealing it
External actionSafety and permissions evaluationThe agent does not execute sensitive actions without defined controls
07

A procurement and evaluation protocol prevents excessive promises

Anyone receiving a BrowseComp figure from a vendor should first request the experimental sheet. At a minimum, it should include the dataset version, execution date, exact model, search and browsing tools, budgets, number of trajectories, aggregation system, and grading criterion. It should also state how many items were not evaluated and how broken links, blocks, or infrastructure errors were counted. Without these details, the percentage has limited meaning and comparison with another figure is fragile.

The next step is to reproduce the relevant capability with an internal test. It is advisable to build a small but representative set of questions that the product needs to solve, without publishing their answers while it remains an active evaluation. It should include permitted real documents, expected access restrictions, ambiguous queries, recent information, and, where applicable, cases with contradictory sources. The aim is not to crown a single number, but to observe failure modes and decide on controls.

The internal evaluation can separate phases. First, measure retrieval: does the agent find relevant evidence? Then measure justification: can it explain which document supports each conclusion? Finally, measure decision or action: does it abstain, request review, or escalate appropriately when evidence is weak? This separation prevents strong search performance from concealing poor behavior in higher-risk tasks.

For comparisons among systems such as Claude Sonnet 5, Claude Fable 5.1, or other agents, the same rule applies: do not infer capability differences from isolated percentages if the harness does not match. A useful comparison requires running equivalent configurations or, when that is not possible, explicitly describing the differences. A commercial name, a model card, or an advertised figure does not replace that experimental control.

Checklist before deploying a browsing agent

  1. 01Request the complete protocol accompanying any external BrowseComp result.
  2. 02Check whether the product task ends in a brief fact or requires synthesis, citations, freshness, or action.
  3. 03Design an internal dataset with sources and constraints similar to those of the real environment.
  4. 04Measure retrieval, evidence quality, ambiguity handling, and action safety separately.
  5. 05Define abstention thresholds, human escalation, and trace logging before deployment.
  6. 06Re-evaluate periodically if the product depends on the live web, search engines, or changing sources.
08

Conclusion: narrow, valuable, and insufficient evidence

BrowseComp provides a useful measurement of a capability that is often difficult to observe: finding a specific fact when evidence is dispersed and browsing requires persistence. Its format of short, verifiable answers enables relatively direct evaluation across many problems. For teams building or buying search agents, ignoring that signal would mean losing relevant information.

A rigorous interpretation requires maintaining the scope of the claim. The benchmark does not turn an accuracy rate into a guarantee of reliable research, nor does it automatically demonstrate source quality, explanation, currency, ambiguity resolution, neutrality, or operational safety. In addition, because it runs in a changing web environment, a figure needs a date, tools, and conditions to be intelligible.

The practical conclusion is not to dismiss BrowseComp, but to use it as one piece of evidence within a broader evaluation. A vendor should be able to describe its harness; a buyer should be able to repeat a test fitted to its use case; and a product team should retain controls for cases where the web, the sources, or the consequences of an answer make a correct brief fact insufficient. In that way, the benchmark serves the purpose for which it was designed without inflating its promise.

Open questions

  • The supplied information does not specify the exact behavior of the current official evaluator for spelling variants, aliases, answer normalization, or manual review; that rule should be confirmed in the current implementation before reproducing results.
  • No complete configurations are provided for specific model or vendor figures, so it is not possible to attribute or compare particular results across models.
  • The availability of pages and search results may have changed since the runs described in the sources; a repetition on the open web may produce different results.
  • No evidence is supplied that a higher BrowseComp score quantitatively predicts performance on the specific distribution of queries relevant to each organization.
  • The exact operational relationship between BrowseComp-Plus and BrowseComp should be verified by task, corpus, and protocol, rather than by name alone.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction