From a fixed list to an adaptive search
CART—short for Closed-Loop Adaptive Red Teaming—is an evaluation framework that proposes adapting safety tests as results emerge. Its central idea is that a finding should not merely be a response recorded at the end of a campaign: it can also help guide the next attempt. The work is presented in a preprint, which should not be mistaken for independent validation or a guarantee of safety.
The problem it aims to address is specific. A fixed collection of prompts can help check known risks and make repeatable comparisons possible, but it does not necessarily change when a test uncovers an unexpected weakness. According to the study summary, CART begins with broad risk coverage, continues exploring weaknesses that emerge, and aims to ensure that new tests do not become repeated variations of the same attack.
This approach does not mean static tests are no longer useful. A fixed test suite can provide a shared baseline; the adaptive approach adds a search guided by what happens during the evaluation. The question is whether this adaptation finds relevant failures that replaying the initial cases would not have detected, and under what conditions. The summary states that CART discovers more failures and higher average risk than static replay of seed tests for every Target with an available baseline. The material summarized here does not provide figures with which to measure the size of that difference.
The comparison, therefore, concerns what two testing policies discover within the evaluations conducted. By itself, it does not establish that the systems tested fail more often in everyday use, or that CART has covered every possible risk.
What the approach compares
The methodological difference is whether the next test case adapts. The table does not imply that one strategy replaces the other in every use case.
| Approach | What it does | What it can reveal | Limitation |
|---|---|---|---|
| Static replay | Runs an initial set of tests again. | Results for known cases and a repeatable baseline. | It does not necessarily explore new weaknesses that emerge during a campaign. |
| CART | Uses evaluation results to guide later tests while aiming for broad, diverse exploration. | Whether adaptation discovers failures that do not appear when seed tests are replayed. | Results depend on the tests, roles, and targets that were evaluated. |
Three separate roles in the cycle
The framework distinguishes three roles. The Challenger proposes tests; the Target is the model or agent that receives those tests; and the Judge evaluates the results. Separating these roles makes it possible to study each function independently and avoids describing CART as a single model that decides, responds, and verifies all at once. The Target may be a text-only model or a tool-using agent, although the available summary does not identify the specific models or detail the tools used.
In an adaptive cycle, the Target’s result can provide signals for deciding which vulnerability to explore next. The Challenger generates new tests guided by those signals; the Judge assesses the responses; and the process preserves evidence and the provenance of findings. The summary specifically highlights recording the evidence and origin of each finding, but the material available here does not specify the precise fields included in that record or its format.
Separating the roles does not eliminate the risk of error. A Judge may misclassify a response, and the choice of Challenger and Judge can affect which evidence emerges. CART’s own summary says that Challenger–Judge choices affect the evidence discovered. For that reason, the number of reported failures should not be treated as independent of the decisions made to generate and assess them.
The evaluation cycle described by CART
- 01Begin with tests that cover a broad set of risks.
- 02Send a test to the Target, which may be a text-only model or a tool-using agent operating within defined limits.
- 03Evaluate the response with the Judge and preserve the evidence and provenance associated with the finding.
- 04Use the results to guide new tests, aiming to broaden exploration without repeating the same failure.
- 05Review which Challenger and Judge combination produced the evidence before interpreting the result.
What results the preprint reports
The summary groups the experiments into three families: Frontier, JAH, and Agentic. It says that, for every Target with an available baseline, CART found more failures and higher average risk than static replay of seed tests. It also states that the gains extend to tool-mediated agent tests. According to the experiments described, this suggests that contextual adaptation can uncover weaknesses that direct prompt replay does not exercise.
There are important limits to this interpretation. The summary does not provide here the number of Targets evaluated, model names, detection rates, numerical differences between methods, or details of each experimental family. Nor does the word “more” alone establish the statistical or practical significance of the difference. Assessing those aspects requires consulting the complete results and their conditions, rather than inferring them from the summary.
The authors also caution that the results describe what the testing policies discover, not how often failures occur in real-world deployments. A red-teaming campaign selects inputs, conditions, and evaluation criteria; it is not equivalent to observing every interaction a system would have in production. CART can therefore provide evidence about the performance of a search strategy without directly estimating the incidence of real-world failures.
What can be concluded—and what remains open
| Question | What the summary indicates | What it cannot establish on its own |
|---|---|---|
| Which evaluation families are described? | Frontier, JAH, and Agentic. | The full details of each protocol. |
| How does CART compare with the baseline? | It reports more failures and higher average risk for every Target with an available baseline. | The numerical size of the difference and its practical significance. |
| Were agents tested? | The summary reports results from tool-mediated agent tests. | Which agents and tools were used, and what specific limits applied. |
| How often do systems fail in production? | The work clarifies that it measures what the testing policies discover. | The frequency of failures in real-world deployments. |
Traceability, diversity, and outstanding limits
Recording the provenance and evidence for a finding matters because it makes it easier to review how the finding was produced and what observation supports its classification. Without that information, a list of results can be difficult to interpret or reproduce. CART says that it records evidence and origin, but the summary does not describe the data schema, identifiers, availability of the records, or whether an independent person can reconstruct each test.
Diversity also needs an operational definition. Saying that tests are intended to be diverse does not demonstrate that the cases differ in the relevant sense: their wording may change while they target the same vulnerability, or they may appear similar while exploring different mechanisms. The summary says that CART keeps new tests diverse, but it does not provide the metric, threshold, or procedure used to verify this. It is therefore not possible to confirm from this description alone how duplicates are identified or how superficial variations are prevented from being rewarded.
Reproducibility is also unresolved. Repeating an evaluation often depends on factors such as model versions, instructions, enabled tools, search budget, evaluation harness, and scoring rules. The material provided does not confirm which CART components have been published, or whether the code, generated tests, and findings are available for an external replication. That availability should not be presented as fact without verification.
In real-world deployments, interpretation also depends on the product and its configuration. A result obtained through an experimental endpoint does not necessarily describe the behavior of a production interface or configuration. An evaluation should make clear what was tested and under which limits, and consider additional testing of the system that is actually deployed. This does not change the results reported by CART, but it does constrain the decisions they can support.
An evaluation contribution, not a safety certificate
CART’s proposed contribution is methodological: turning the results of a campaign into information that can guide the next round of tests, rather than relying only on a list that stays fixed. The preprint reports advantages over a static baseline across three evaluation families, including tests with agents and tools. That evidence makes the approach worth examining, but it is not enough to claim that CART detects every risk, works equally well for every model, or by itself improves a product’s safety.
To decide how much confidence to place in the proposal, the practical questions are: which Targets were compared, what conditions and budgets applied, how diversity was verified, how Judge errors were controlled, and whether third parties can reproduce the results. It is also worth checking whether the tests match the model and endpoint that will actually be used. Until verifiable answers are available, CART should be read as a promising adaptive-search framework with results reported by its authors and acknowledged limitations—not as a guarantee of safety.
Criteria for interpreting an adaptive evaluation
- 01Identify which models, agents, tools, and configurations were tested.
- 02Compare the adaptive strategy with a clearly defined baseline and examine the size of the differences, not just the direction of the results.
- 03Check how test diversity and deduplication were measured.
- 04Review the human or independent checks applied to automated judgments.
- 05Confirm which evidence, code, and records are available for replication.
- 06Do not extrapolate results from a campaign to the frequency of failures in production without specific evidence.
Open questions
- The material provided does not identify the models and agents evaluated, their versions, or the limits placed on tools.
- Detailed figures, sample sizes, and uncertainty estimates are not provided here to quantify the advantage over the baseline.
- The summary does not specify how test diversity is measured or how attacks exploring the same failure are deduplicated.
- It does not confirm which independent or human mechanisms are used to detect false positives and Judge biases.
- It does not confirm whether the code, generated tests, and records are available for external replication.
- The preprint results do not establish the frequency of failures in systems deployed under real-world conditions.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction