Ilustración editorial para CAVEAT mide si los agentes de compra mantienen el objetivo del usuario ante incentivos de plataforma
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

A conflict of interest distinct from a malicious instruction

Computer-use agents can act on a person’s behalf across websites and apps: searching for options, comparing information and, in some cases, completing a purchase. CAVEAT asks a specific question about this delegated use: does an agent stick to the user’s priorities when the environment it operates in has its own incentives to favor a different outcome?

The distinction matters. In an explicit attack, someone may try to insert deceptive instructions to make a system disobey the user. CAVEAT instead focuses on steering mechanisms attributed to the environment. The concern is not necessarily that a malicious command appears, but that the way a commercial situation is presented or structured could sway the agent’s decisions. The authors’ abstract describes this possibility as an alignment problem between platform incentives and the user’s goal.

That makes this a question of robustness to incentives, not a general test of whether agents are safe or unsafe. The benchmark asks whether, under the specific conditions defined for its tests, an agent chooses the product the authors consider optimal for the user. On its own, it does not establish what happens across all delegated purchases or what motivates any particular platform.

Two different questions

Type of riskWhat is being testedWhat it does not establish on its own
Explicit malicious instructionWhether an agent resists instructions intended to divert it from the user’s request.Whether it preserves the user’s priorities when the environment steers it without issuing an explicit instruction.
Misaligned incentives in the environmentWhether an agent sticks to the user’s goal when faced with steering mechanisms defined in the benchmark.That a real store manipulated a purchase, or that all agents fail in the same way.
02

Nine environments and eight mechanisms, with details still to verify

According to the paper’s abstract, CAVEAT brings together nine marketplace environments and evaluates them using a taxonomy of eight common steering mechanisms. The benchmark is intended to address a gap the authors identify in earlier evaluations: testing whether agents preserve user goals when the environment itself has a stake in the outcome, rather than focusing only on cooperative settings or explicit attacks.

That scope is relevant, but the supplied abstract does not list the nine environments or describe the eight mechanisms one by one. It is therefore not possible to say here exactly which elements of a page, messages or interaction rules were changed in each case. Nor can we judge from the high-level description alone whether the mechanisms represent common practices, extreme cases or a mixture of both.

The word “controlled” helps explain the benchmark’s purpose: comparing results under conditions prepared for evaluation. It does not mean that those conditions have been shown to reproduce how real digital retailers operate. Assessing that correspondence would require details about the tasks, interfaces, data and criteria used to construct each environment.

What is known and what remains to be verified

ElementWhat the abstract reportsDetails not confirmed by the available material
EnvironmentsNine marketplace environments.Their names, characteristics and degree of similarity to real stores.
MechanismsA taxonomy of eight common steering mechanisms.An individual description of each mechanism and its frequency or realism.
EvaluationA comparison between matched control episodes and episodes with mechanisms enabled.Full instructions, tasks, setup and scoring rules.
ReproducibilityThe abstract reports results and an intervention called CAVEAT-Harness.Whether CAVEAT’s code, prompts, environments and data are available.
03

The observed drop is large, but it belongs to the test

The central result reported in the abstract compares matched control episodes with episodes in which steering mechanisms are enabled. In the control episodes, agents bought the product deemed optimal for the user in 78.6% of cases; in the episodes with steering, they did so in 17.3%. The authors say they evaluated five model families. The abstract does not name those families or break down results for each one.

The difference supports a bounded conclusion: under the conditions evaluated, the mechanisms included in CAVEAT coincided with a substantial decrease in purchases classified as optimal for the user. It is not enough to conclude that agents used by the public have a 17.3% success rate in real purchases. That figure describes benchmark episodes, not a representative sample of transactions in external marketplaces.

The authors also report that larger models and increased reasoning effort improve robustness, although substantial failures remain. This statement does not identify which specific model performs best or how much each change contributes, because the abstract gives no model-by-model figures, uncertainty intervals or details of the comparisons.

The authors’ trajectory analysis identifies three points where steering can enter the decision process: the agent distorts the user’s priorities, narrows the alternatives it considers too early, or makes a decision before resolving relevant information. These are diagnoses reported by the paper. The available abstract does not provide the examples behind each category or quantify how much each contributes to the overall result.

How to interpret the comparison

The comparison reports performance within CAVEAT’s design. It is not a direct estimate of how often errors occur in everyday use.

  1. 01In matched control episodes, the authors record whether the agent buys the product defined as optimal for the user.
  2. 02They repeat the evaluation with steering mechanisms enabled.
  3. 03They compare the proportions: 78.6% in the control condition and 17.3% with steering, according to the abstract.
  4. 04They interpret the difference as evidence of vulnerability under those conditions, not as a universal failure rate.
04

CAVEAT-Harness improves the result, according to its authors

In response to the failure points they identify, the authors develop CAVEAT-Harness, an intervention aimed at the three stages described: preserving the user’s priorities, keeping alternatives open and resolving relevant information before deciding. The abstract says this tool raises user-optimal purchases by 55.0%.

That figure should be retained as the authors present it. The abstract does not clarify whether 55.0% means a relative increase, a difference in percentage points or another calculation. The supplied material also does not specify the exact comparison setup or whether the improvement holds equally across all nine environments and five model families. Without those details, it would not be appropriate to convert the figure into another metric or attribute broader reach to it.

The paper also reports that post-training further improves a smaller open model. This is a promising research direction, but the abstract does not say which model was tuned, what data were used or how much it improved. It therefore does not allow the procedure’s reproducibility or cost to be assessed.

05

What CAVEAT does not show about real purchases

The main caveat is the distance between a controlled benchmark and a live commercial platform. The available abstract does not say whether the evaluated interfaces reproduce existing stores, whether tasks were tested with human participants, or whether user preferences were expressed in a way comparable to a real purchase. Without that information, it is not possible to measure how far the results generalize.

Nor can we claim that a platform deliberately directed an agent toward a particular product. The evaluation is intended to test robustness to misaligned incentives, but the abstract presents no evidence about the practices of real companies. An experimental scenario that models a risk and an allegation about a company are different kinds of claims.

Related research examines a different question. SusBench, for example, is described as an online benchmark of computer-use agents’ susceptibility to dark patterns, evaluated across 55 consumer websites and involving human participants. That approach may offer a point of comparison for consumer interfaces, but it is not equivalent to CAVEAT: the approaches and evaluation questions are not identical. The existence of that work also does not automatically validate CAVEAT’s results.

The supplied sources do not make it possible to confirm the full contents of the CAVEAT paper, whether its reproducibility materials have been published, or the limitations its authors discuss beyond the summary. It also remains unclear which specific models took part. These gaps do not invalidate the reported figures, but they limit how much can be independently examined and which conclusions are warranted.

What the results allow us to say

ConclusionStatus
In the reported episodes, the share of user-optimal purchases was lower with steering mechanisms than in the control condition.Supported by the figures in CAVEAT’s abstract.
The evaluated agents fail in 82.7% of real purchases.Not supported: 17.3% refers to benchmark episodes.
A specific commercial platform is manipulating agents.Not demonstrated by the reported results.
CAVEAT-Harness guarantees correct decisions in new environments.Not demonstrated: an improvement in the evaluation is reported, not a general guarantee.
06

What it would take to evaluate shopping agents

Before using a benchmark as a basis for certifying shopping agents, its tasks, environments and operational definition of the “optimal product” would need to be available for review. Results would also need to be published separately by environment and model, with an explanation of how improvements were calculated and how measurement variability was quantified. CAVEAT’s abstract does not provide all of these details.

A fuller evaluation should test whether results hold across different interfaces and varied tasks, and when user preferences are expressed in more than one way. It would be useful to compare controlled scenarios with tests in representative commercial environments, while protecting personal data and defining independent evaluation criteria. It would also be important to measure whether interventions improve adherence to the user’s goal without introducing new errors, such as unnecessary delays or decisions based on insufficient information.

Making code, instructions, environments and data available would let other teams reproduce the tests and check the trajectory analyses. The supplied material does not confirm that these resources have been published for CAVEAT. Until that is clarified, reproducibility remains an open question, not a feature that can be taken for granted.

According to its abstract, CAVEAT offers a way to examine a problem that goes beyond explicit attacks: the decision-making context may have incentives that differ from the user’s. Its figures justify further study of that possibility and testing targeted defenses. On their own, however, they are not enough to characterize the behavior of all agents, accuse a real platform or certify a shopping system. For readers following broader news, comparisons and developments in the sector, that is the key distinction: the benchmark provides an experimental signal from constructed scenarios, while its impact on everyday use still requires additional evidence.

Checklist before certification

  1. 01Publish a verifiable definition of the user’s priorities and how the optimal option is identified.
  2. 02Document the tasks, environments, steering mechanisms and configuration of each agent.
  3. 03Report results by model and environment, alongside statistical uncertainty and comparison criteria.
  4. 04Allow third parties to reproduce the evaluation using relevant code, prompts, data and materials.
  5. 05Check transfer to representative interfaces and situations without confusing test scenarios with evidence of business conduct.
  6. 06Evaluate interventions against mechanisms not used during their development and examine possible side effects.

Open questions

  • The supplied material does not list the nine environments or individually detail the eight steering mechanisms.
  • The model names, model-by-model results and full definition of the user-optimal product are not provided.
  • It is unclear whether the 55.0% increase attributed to CAVEAT-Harness is relative or expressed in percentage points.
  • It is not confirmed here whether CAVEAT’s code, prompts, environments and data are available to reproduce the evaluation.
  • There is insufficient information to determine how closely the benchmark’s interfaces and tasks represent real commercial platforms.
  • The abstract does not establish whether the conclusions hold across a broader range of models, tasks and shopping scenarios.
07

Keep exploring

07

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction