Completing a task does not prove that every visual query was necessary
Visual agents can request an enlarged or cropped image while working on a task—for example, to read a detail that is hard to make out in the initial view. If they then answer correctly, an evaluation based only on the outcome may treat the call as a success. But that leaves two questions unanswered: did the agent need that information to solve the task, and, once it received the information, did it actually use it?
A preprint titled “When Should a VLM Look? Paying Only for Visual Calls That Were Needed and Used” argues that these questions matter because a successful call does not, by itself, prove either that the call was necessary or that the agent used the pixels it obtained. The authors introduce CounterCredit, a training method that aims to evaluate both conditions for every call that returns an image.
The distinction matters when interpreting an agent’s scores. A system can answer a question correctly and still have made a redundant query; it can also request a visual region without basing its answer on that region. In either case, simply rewarding the call because the final answer was correct could assign it credit that is not justified. This is a question of evaluation and training design, not proof that all current visual agents behave this way.
How CounterCredit separates necessity from use
According to the preprint, CounterCredit makes both checks at the specific state immediately before each call. To estimate necessity, it compares the branch in which the agent queries the image with the option of answering immediately, without making the query. To estimate whether the agent used the evidence, it compares the returned image with random patches of the same size, substituted for that image in the same call. The approach is intended to distinguish the value of looking from the value of the specific visual information received.
The system uses the policy’s own score for the correct answer, described in the abstract as a gold-answer score. If a call passes both checks, it earns a cashback reward; other executed calls incur a cost, which the paper calls “rent.” The authors say they cap this cost so that every correct trajectory still ranks above every incorrect trajectory. They also describe a dual-channel GRPO advantage that keeps the cost in its own units.
These operations are criteria in the proposed method, not direct observations of the model’s internal intent. Comparing against an immediate answer and substituting random patches are tests defined by the authors. Their validity therefore depends on whether those comparisons adequately reflect what the method is meant to measure.
The two checks applied to a call
- 01Record the agent’s state immediately before the visual call.
- 02Check necessity: compare the branch that queries the image with the option of answering without querying it.
- 03Check use: replace the returned image with random patches of the same size and compare the score.
- 04Issue cashback only if both necessity and use are verified; otherwise, apply the cost specified by the method.
What figures the preprint reports, and which tests they come from
The preprint’s abstract reports several results, which should be kept tied to the evaluation sets on which they were measured. For the cold-start checkpoint, the authors say that only 10% to 12% of visual calls were both needed and used. For released agents, they report spurious-call rates ranging from 36% to 87% on individual benchmarks. The abstract also says that about two thirds of what an outcome-based system rewards goes to calls that were neither needed nor used.
Starting from the same checkpoint, prompt pool, and budget, CounterCredit reaches 89.5% on V*, 80.2% on HR-Bench-4K, and 76.4% on HR-Bench-8K. The authors compare these figures with outcome-only GRPO and report an advantage of 6.3 to 9.4 points, with 1.78 calls per question versus 1.84. They also report that the method brings the spurious-call rate down to 31% to 36%, the lowest among the agents they evaluated.
The paper further reports that the same recipe raises the average score of a Qwen3-VL-8B base model from 75.4 to 80.8. These figures are results reported by the authors in the preprint. They are not a universal measure of agent quality, and on their own they do not show that the method improves performance on tasks outside the evaluation.
Results reported in the preprint abstract
| Test or comparison | Reported result | Scope |
|---|---|---|
| Cold-start checkpoint | 10% to 12% of calls were both needed and used | Calls made by the evaluated checkpoint |
| Released agents | 36% to 87% spurious calls | Varies by individual benchmark |
| V* | 89.5% | Score reported for the benchmark |
| HR-Bench-4K and HR-Bench-8K | 80.2% and 76.4% | Scores reported for each benchmark |
| Calls per question | 1.78 with CounterCredit versus 1.84 with outcome-reward GRPO | Comparison described by the authors |
| Qwen3-VL-8B base | Average score increases from 75.4 to 80.8 | Result reported for the applied recipe |
What the benchmarks do—and do not—tell us
V* is a benchmark related to guided visual search; the paper that introduced it helps trace the benchmark’s origins. HR-Bench-4K and HR-Bench-8K are high-resolution image-perception evaluation sets introduced in another paper. CounterCredit’s scores on these benchmarks describe its performance on those specific sets, but do not automatically demonstrate the same behavior in browsing, everyday visual assistance, or other tasks.
There is another point to keep in mind: the spurious-call rate and the final score answer different questions. A lower rate suggests that, by the study’s criteria, fewer calls failed to meet both requirements. It does not, by itself, prove that answers are safer, that the system makes better use of every kind of image, or that the total cost is lower under all conditions. Nor should the cold-start checkpoint results, released-agent results, and Qwen3-VL-8B experiment be conflated: the authors describe them as separate comparisons.
A preprint, with open questions about replication and generalization
The work is presented as a preprint, and the supplied listing marks it as under review. Its figures should therefore be treated as the authors’ initial results, not as conclusions established through completed peer review. The available documentation supports the description of the method and the figures in the abstract, but is not enough to determine independently whether every evaluation choice remains robust across other models, images, or tasks.
Nor can we conclude from the sources available for this article whether enough code and data are publicly available to reproduce the full set of experiments, or whether independent groups have repeated the results. The fact that this information is absent from the sources consulted does not prove that such resources or tests do not exist; it means we should not claim they do without a source that verifies it.
To find out whether CounterCredit generalizes, useful tests would include independent comparisons with other models and policies, additional benchmarks, and experiments with different budgets and tools. It would also help to publish the details needed to reproduce the evaluation decisions, compare different versions of the necessity and use criteria, and report results by task rather than only as averages. Such tests could clarify whether the method reduces redundant calls without penalizing queries that matter in other scenarios.
For now, the most measured conclusion is limited: the preprint presents an explicit way to reward calls it considers necessary and used, and reports improvements on the benchmarks studied compared with an outcome-only baseline. Whether those improvements extend beyond the evaluated settings remains an open question.
Open questions
- The supplied sources do not confirm whether the code and data needed to reproduce all experiments are publicly available.
- The supplied sources do not establish that independent evaluations have replicated the results.
- It is not known whether the metrics and improvements hold with other models, tools, budgets, or tasks beyond the cited benchmarks.
- The review status may change; the listing for the main source identifies it as a preprint under review.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction