What it means to outperform a human reference on a bounded test
A benchmark score answers a narrowly defined question: how a model performed on a particular dataset, for a particular task, under particular instructions and conditions. It does not automatically answer a much broader question: whether the model can correctly read a lab’s protocols, interpret its data, and reliably support scientific conclusions. To decide whether Claude Sonnet 4.5 could help in a real workflow, those two claims need to remain separate.
The Claude Sonnet 4.5 system card reports LAB-Bench results for several tasks, including Protocol QA, FigQA, SeqQA, and Cloning, and includes a human reference. In the graph identified for Protocol QA, Sonnet 4.5 scores 0.833 at k=10, compared with 0.741 for Sonnet 4. These are figures from Anthropic’s published result, not an independent validation of performance on any laboratory’s documentation.
So a score above a human reference on a specific evaluation does not mean the model is better than researchers in general, gets every question right, or produces answers safe to turn into experimental instructions. Interpretation depends on the questions asked, how they were scored, and who was included in the human comparison group. The available source information is not enough to transfer the result to other documents, teams, or decisions.
Identify the model that was actually evaluated
“Claude Sonnet 4.5” names a commercial family or version, but a reproducible evaluation needs to record the identifier sent with the request. Anthropic’s documentation distinguishes dated identifiers from aliases and identifies Sonnet 4.5’s snapshot as claude-sonnet-4-5-20250929. Recording that identifier, along with the test date and configuration, reduces ambiguity when comparing results.
This precaution also matters when comparing later models. The supplied official source for Sonnet 5 announces that model, but does not establish that it was evaluated on the same LAB-Bench or BixBench tasks, using the same protocol and dataset as Sonnet 4.5. The supplied documentation also does not allow us to claim that an equivalent comparison exists for Sonnet 4.6. The release of a later version is not, by itself, evidence of better performance on these tasks.
A useful comparison needs to hold the task and conditions in view, rather than simply set model names against each other. If the questions, number of examples provided in the prompt, available tools, or scoring method change, those differences may contribute to the result alongside the model version.
What to record before comparing versions
This table is a checklist for an internal test; it does not attribute methodological details to published evaluations when the available sources do not specify them.
| Item | Minimum record | Why it matters |
|---|---|---|
| Model | Exact identifier and date of access | Distinguishes snapshots from aliases. |
| Task and corpus | Version, partition, and exclusion criteria | Helps prevent comparisons between different questions or data. |
| Configuration | Prompt, examples, tools, and parameters used | Makes it possible to interpret the conditions that produced the response. |
| Evaluation | Rubric, reference answers, and reviewers | Makes explicit how an answer was judged correct. |
Three benchmarks, three kinds of evidence
LAB-Bench is a collection of tasks for measuring language-model capabilities related to biological research. The original paper describes multiple-choice questions and a comparison with expert researchers. Its scope is the tasks included in the benchmark; it is not a direct observation of laboratory work or an evaluation of every stage of research.
Protocol QA, part of LAB-Bench, focuses on answering questions about protocols. The published result for Sonnet 4.5 provides evidence about its performance on that task under the reported evaluation conditions. It does not show that the model can carry out a protocol, detect every ambiguity, or correctly adapt an instruction to local reagents, instruments, and procedures. Understanding a question about documentation and performing an experimental operation are different capabilities.
LAB-Bench also includes tasks associated with reading figures, sequences, and cloning. System-card results in these categories broaden the evaluation map, but the scores are not interchangeable: each task tests different capabilities and formats. A strong result on protocol questions does not automatically imply an equivalent result when interpreting a figure.
BixBench addresses computational biology through open-ended questions that require multistep analytical trajectories. The original paper describes initial evaluations with GPT-4o and Claude 3.5 Sonnet, not Claude Sonnet 4.5. BixBench is therefore useful for explaining one kind of computational-analysis evaluation, but its cited initial results are not Sonnet 4.5 results and should not be presented as such.
The BixBench repository describes a harness and configuration for running evaluations that include notebook and code analysis. This structure makes it possible to plan a controlled reproduction of the benchmark, but it does not remove the need to check which model version, tools, and conditions were used in each experiment. For LAB-Bench, the authors’ repository documents evaluation formats and scripts, and notes that the public materials do not contain the complete dataset. This matters if the goal is to reconstruct a published score exactly.
What each result can and cannot support
This distinction helps prevent a benchmark task from becoming a broad claim about scientific competence.
| Evaluation | Evidence it provides | What it does not establish on its own |
|---|---|---|
| Protocol QA | Performance on protocol questions for the dataset and configuration evaluated. | Correct execution, adaptation to local procedures, or experimental safety. |
| FigQA and other LAB-Bench tasks | Performance on specific biological research tasks included in LAB-Bench. | Uniform ability to analyze every figure, sequence, or biological problem. |
| BixBench | Evaluation of open-ended computational biology tasks involving multistep analysis. | A Sonnet 4.5 result when the cited evaluation concerns other models. |
Why Protocol QA scores should not be combined without checking
The supplied evidence includes a Protocol QA result in the Sonnet 4.5 system card: 0.833 for Sonnet 4.5 at k=10 and 0.741 for Sonnet 4. The editorial proposal notes that some Anthropic publications show different scores for tasks called Protocol QA. However, the supplied sources do not document enough to determine whether those figures come from the same dataset, different partitions, different prompts, or other methodological changes.
Having the same task name does not guarantee that two measurements are equivalent. Before presenting them as a change over time or a discrepancy, one would need to check the exact model version, evaluation dataset, question format, examples included in the prompt, tool availability, scoring rule, and handling of partial answers. The meaning of k=10 in the graph and how responses were aggregated would also need to be clarified before interpreting that parameter precisely.
The human comparison requires a similar audit. The system card shows a human reference, but the materials described here do not provide enough detail to confidently characterize the population, number of participants, their expertise, or the instructions they received. Without those details, it would not be rigorous to present that reference as a universal standard of expert performance.
The appropriate conclusion is not that the results are incompatible, or that one figure is wrong. It is that comparability cannot be resolved from the available information. A rigorous account reports each figure in context and makes the uncertainty explicit, rather than calculating a trend or combining scores.
From benchmark to lab: the gap that remains
A benchmark can test whether an answer matches an answer key, but a lab needs to know whether that answer is correct for its current documentation and the context of the question. Protocols may have local variants, internal names, revision notes, or relevant conditions that do not appear in the reference material for a public test. The ability to answer general questions does not verify that the model found and applied the relevant version.
Answering a question about a protocol is also not the same as carrying it out. Execution depends on people, equipment, samples, materials, and controls—none of which are demonstrated by a question-and-answer score. Likewise, producing a plausible interpretation of a figure does not prove that the interpretation is valid or that the scientific conclusion holds up against the full data.
Tool use adds another dimension. In a computational-analysis task, the ability to produce or execute code may affect the result; in a documentation task, the model may respond only from the prompt. To attribute a difference to a model version, the comparison needs to hold tools constant and record when they were used. The available sources describe elements of benchmark evaluation, but are not sufficient to reconstruct all conditions behind every published figure.
For these reasons, a favorable result is a reason to evaluate a bounded use—for example, finding a candidate answer in documentation and preparing a draft for review—not permission to replace specialist verification or proof of scientific validity.
How to design a useful retrospective validation
An internal test does not need to automate experiments or expose sensitive information to evaluate document comprehension. It can be built from an authorized, frozen corpus: documents the team is permitted to use, an identifiable version of each file, and a record of what material was available when the model answered. Freezing the corpus prevents later document changes from altering the interpretation of results.
Specialists familiar with the documents should write and review representative questions, along with reference answers verified against the source. It is useful to include questions with explicit answers, questions that require combining information from multiple parts of a document, and questions that cannot be answered from the available material. This last category tests whether the system recognizes the limits of the evidence instead of filling gaps with a plausible response.
Each answer should be evaluated against a rubric before models are compared. A rubric can separate content accuracy, support in the source, identification of the relevant version, important omissions, and unsupported claims. People judging the results should not depend on the model’s brand or version, and disagreements between reviewers should be recorded and resolved using previously agreed criteria.
To compare Sonnet 4.5 with later versions, use the same questions, documents, instructions, tools, and scoring rules. Record the exact identifier for every model. If a platform does not allow a version or parameter to be fixed, that limitation is part of the result and should not be hidden. Repeated runs can also help identify whether the system responds consistently; the number of repetitions should be chosen according to the cost and risk of the application, not inferred from public sources.
The test should be retrospective and limited to reviewable document comprehension. There is no need to ask the model to design, modify, or carry out experiments to find out whether it can locate information, interpret text, or recognize that a question is unanswered. If figures or analytical results are evaluated, the team needs verified references and should keep those results separate from protocol-question scores.
Document-evaluation protocol
A suggested design for measuring a bounded task without confusing it with experimental execution.
- 01Agree on permitted use and select authorized documents; record their versions and freeze the corpus.
- 02Have specialists write questions and verify every reference answer against the document.
- 03Include answerable cases, cases requiring integrated information, and cases without a sufficient answer.
- 04Set prompts, tools, models, scoring criteria, and number of repetitions before running the test.
- 05Review answers blind, record disagreements, and classify errors using a shared rubric.
- 06Compare results by category and document limitations, serious failures, and conditions of use.
Classify the errors that matter
A single overall accuracy rate can hide failures of very different severity. In protocol reading, an omission may leave out an important condition; reversing steps may alter the meaning of an instruction; confusing units may make an answer look precise when it is not. The team should record these cases separately rather than assume that every incorrect answer carries the same risk.
For figure questions, distinguish describing a visible element from claiming a relationship that is not shown or proposing a causal explanation. In computational tasks, distinguish reproducible analysis from a conclusion that does not follow from the data. Also flag references that do not support the answer, citations to the wrong sections, and unwarranted confidence when information is missing.
Editorial safety evaluation should stay within the documentary task: check whether the model faithfully represents authorized content and identifies uncertainty. There is no need to include sensitive experimental instructions or ask for actions outside the reading task. The goal is to determine whether the system can serve as a reviewable aid, not to turn an internal benchmark into a validation of procedures.
A minimum error taxonomy
This classification helps determine whether an aggregate score hides an error category that is unacceptable for the intended use.
| Category | What to review | Example signal |
|---|---|---|
| Omission or reversal | Whether a relevant detail is missing or the stated order is changed. | The answer omits a condition or reverses two steps from the text. |
| Units and conditions | Whether units, limits, and document context are preserved. | A number appears without a unit or is assigned to a different condition. |
| Figure reading | Whether visible observation is separated from interpretation. | The answer presents an explanation as though it were labeled in the figure. |
| Documentary support | Whether the answer is supported by the cited section. | The reference does not contain the attributed claim. |
| Uncertainty | Whether unanswerable questions are recognized as such given the corpus. | The model gives a categorical answer without support. |
Bounded use and conclusion
If a favorable internal evaluation confirms that Sonnet 4.5 can find information and represent it with verifiable support, the team could test it as a reading assistant or as a generator of drafts for review. The final answer would still depend on the lab’s agreed verification process. A public score alone does not justify using an output as an experimental instruction or letting it determine a scientific interpretation.
Comparison with successors should be treated as a new controlled test. The supplied documentation identifies Sonnet 4.5 and points to an official announcement of Sonnet 5, but it does not provide equivalent results for Sonnet 5 on Protocol QA, LAB-Bench, or BixBench. Nor does it establish a controlled comparison with Sonnet 4.6. These sources therefore cannot support an inference about which model performs better on the same tasks.
The verifiable conclusion is limited, but useful: Anthropic publishes favorable Sonnet 4.5 results on specific LAB-Bench tasks, including Protocol QA, while BixBench evaluates a different kind of work and the cited initial results do not concern Sonnet 4.5. These data justify a hypothesis that the model may be useful for reading and analysis; they do not establish reliability in a real lab. The missing evidence needs to come from authorized documents, specialist-reviewed questions, traceable references, and comparisons under equivalent conditions.
Open questions
- The verified sources do not resolve whether the different published Protocol QA scores use the same dataset, format, prompting, tools, and scoring method.
- The supplied information does not precisely characterize the population, participant count, or instructions behind the human reference in the graph.
- No evidence is provided for Sonnet 4.6 or Sonnet 5 evaluations on the same tasks under conditions comparable with Sonnet 4.5.
- The LAB-Bench repository notes that the public portion does not contain the complete dataset, limiting exact reproduction of published results.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction