A stated capability does not prove an end-to-end workflow
Cohere identifies Command A+ by the model ID command-a-plus-05-2026 and describes it as supporting text and image inputs, 48 languages, and agentic capabilities. These details help define what is worth testing, but they do not show that the model can reliably solve a task combining all three. Understanding an image, interpreting a request in another language, and deciding whether to call a tool are distinct parts of a task; adding those capabilities together does not guarantee a correct integrated result.
The operational question is not whether Command A+ has vision, supports multiple languages, or can use tools in the abstract. It is whether, in the specific workflow an organization wants to deploy, it can find evidence in an image, understand what the user is asking, and take an appropriate action without inventing information or exceeding its permissions. Answering that requires an evaluation of your own, using representative tasks and observable outcomes.
The provider’s documentation helps establish which model, inputs, and configuration to test. Cohere’s guides describe working with images, tool use, and the Chat endpoint; they should not be mistaken for an independent evaluation of combined performance. Similarly, benchmarks on multilingual multimodal reasoning or visual tool chaining can inform test design, but they do not establish how Command A+ behaves in a particular business workflow.
This approach differs from a comparison between Command A+ and another model for a retrieval-augmented task. The goal here is not to declare a winner or extrapolate from a ranking. It is to evaluate a single model under controlled conditions and decide which tasks it can handle, which need human review, and what configuration or process changes may be needed.
Define a task that requires all three capabilities
The unit of evaluation should be an end-to-end task, not a collection of isolated questions. For example, someone submits a screenshot of an interface and asks, in their own language, whether an operation appears complete and to record the result in a test system. To respond correctly, the model would need to understand the request, identify relevant evidence in the image, decide whether a tool is needed, and, if so, send arguments that match the tool’s schema.
The test should separate what can be observed from what the system is expected to do. A human annotation of the image can serve as the reference for visual interpretation; an expected intent can be used to assess understanding of the request; the tool name and expected arguments can be used to assess the call; and the outcome recorded by the simulated environment can be used to assess the final result. This avoids counting a fluent response as a success when, for example, it misread a number or performed an action using the wrong value.
Before generating examples, define the limits of what the model may do. A simulated tool can return information or perform a reversible operation in a test environment. Do not connect initial tests to production accounts, documents, or processes. Simulation lets you observe tool choice and arguments without confusing model quality with damage or real-world consequences.
It is also worth defining in advance what counts as an appropriate abstention. If the image is illegible, an essential detail is missing, or the request does not authorize an action, asking for clarification may be preferable to executing it. Do not automatically score abstention as failure: its value depends on whether there was enough evidence and on the consequences of acting under uncertainty.
Minimum preparation for each test case
- 01Define the intent, the visual evidence needed, and the permitted action.
- 02Save the test image and note which parts support the expected response.
- 03Specify whether the tool should be used, should not be used, or whether there is not enough information to decide.
- 04Record the expected outcome, including valid arguments and conditions that should lead to abstention.
Build a corpus that represents real use
The evaluation set should include documents and interface screenshots that the organization is authorized to use. It is useful to cover a range of formats and image quality: clear and blurry images, small text, partially cropped elements, different layouts, and, where relevant, tabular data or similar fields. These are proposed design variations, not capabilities guaranteed by the documentation.
Each image should have a human-reviewed reference stating what information is actually present, where it appears, and what uncertainties remain. If a number can be read in two ways or a status is indistinct, the annotation should reflect that. Forcing an ambiguous image into a single “correct answer” label would distort measurement and penalize reasonable abstention.
Language sampling should be based on the users and workflows you expect to support. Although Cohere states that the model supports 48 languages, this is not the same as public evidence of uniform performance in every language or on a particular business task. Record the language of the request, the language of the visual content, and any mixing of the two. If the corpus is translated, a proficient reviewer should check that the instruction retains the same meaning and degree of ambiguity.
To reduce contamination between development and evaluation, reserve cases that are not used to adjust prompts or schemas. You can also create controlled variations from templates, but should not treat those variants as fully independent observations. The priority is to represent real decisions: when a tool is needed, when it adds no value, and when the visual evidence is insufficient.
Recommended condition matrix
You do not need to combine every dimension in every cell. The matrix helps identify which comparisons can attribute a change to the image, language, or tool.
| Dimension | Possible conditions | What it helps you observe |
|---|---|---|
| Input | Text; text and image | Whether the image provides useful evidence and whether including it changes the response |
| Language | Team’s usual language; other relevant languages; visual content in another language | Errors in understanding, reading, and transfer between languages |
| Action | Response without a tool; permitted call; unnecessary call; abstention | Tool selection, decision to act, and appropriate use of evidence |
| Evidence quality | Clear; degraded; incomplete or ambiguous | Robustness and behavior under uncertainty |
Use simulated tools and a reproducible configuration
Each test tool should have an explicit purpose, an argument schema, and controlled responses. For example, a lookup function could return a test status based on an identifier; another function could add a label to a simulated record. Responses should be predictable and saved in the run log. This makes it possible to distinguish a model reading error from an infrastructure failure or an unexpected tool response.
Cohere’s guide to tool use and its Chat reference are starting points for implementing this part. Before running the test, check the current documentation to confirm how tools are declared, which fields the request requires, and which input forms are compatible with the chosen model ID. Do not assume, without checking, that all parameters, formats, or modes available in one configuration can be combined in another.
Keep the configuration constant across conditions except for the factor being measured. Record the exact model ID, date, API or client version, relevant request fields, available tools, and their schemas. If prompts or tool definitions differ between groups, the differences may be due to those changes rather than to the language or image.
Run tests in an isolated environment with reversible effects. Even when a tool is simulated, validate arguments before applying them and retain both the request and the tool response. For a workflow that can modify data, the evaluation should check authorization behavior and how out-of-scope instructions are handled, as well as the text response.
Compare conditions without confusing the causes
The most informative design combines isolated controls with integrated tasks. In a text control, provide the necessary textual information without an image; in another, ask the model to extract information from an image without taking an external action; in a third, evaluate a tool call when the relevant information is already specified. The combined condition requires the model to obtain evidence from the image, interpret the request in the relevant language, and decide what to do.
The controls do not, by themselves, show that the system is ready to operate. They help locate the likely source of a performance drop. If the model extracts a field correctly when it is transcribed but not when it appears in a screenshot, visual interpretation may be the issue. If it understands the field and intent separately but fails when a request in another language is combined with the screenshot, the composition of capabilities warrants attention. If the decision is correct but the arguments are not, the problem lies in preparing the call or in the interface between the model and tool.
For a fair comparison, use the same task and objective across conditions. Keep the expected answer constant and change only the modality or relevant language factor. Vary the order of cases and avoid tuning instructions after inspecting the held-out set. With small samples, report results by case and category rather than presenting an aggregate rate as if it described general performance.
Repeating the same case can reveal variability, but it does not turn a small test into universal evidence. Retain the results of every run, including duplicate calls, incomplete responses, and environment failures. Define success in advance: for example, a task may count as complete only if the evidence is correct, the authorized action was executed with valid arguments, and the final response matches the observed result.
Evaluation sequence
- 01Run isolated controls for visual reading, language understanding, and tool use.
- 02Run combined tasks using the same references and action limits.
- 03Save inputs, final responses, calls, arguments, tool results, and timings.
- 04Manually review a sample of successes, failures, and abstentions before summarizing rates.
- 05Repeat affected cases after correcting the configuration, keeping the adjustment set separate from the acceptance set.
Measure end-to-end success and cost
No single metric adequately describes a multimodal agent. Record separately whether visual evidence was extracted correctly, the request was understood, the appropriate tool was selected, its arguments were valid, execution produced the expected result, and the final response remained faithful to that result. A system can get one stage right and fail at the next; preserving that distinction makes corrective action more specific.
Abstention needs its own measure. Count when the model asks for clarification or says it cannot determine something, and compare that behavior with the actual sufficiency of the evidence. Abstaining when an image is illegible may be appropriate; abstaining when a field is clear and the request is authorized may block the workflow. Similarly, an unnecessary tool call can be a failure even if the tool returns a harmless response.
Record latency and cost per completed task, along with any usage the configuration makes observable. Cohere’s pricing documentation explains the billing model and billable units, but the applicable cost must be checked for the current channel and rate. An average per request can be misleading if some failed tasks require retries or human review. Therefore, also calculate the cost of correctly completed tasks and account for follow-up work.
Do not combine all languages, images, and tool types into a single figure without breakdowns. A high average can conceal a category in which the system misreads identifiers or acts without evidence. Report results by language, image type, evidence quality, tool decision, and error class. If the volume is too small to support stable estimates, describe the results as observations from the tested set, not as a reliable production rate.
Metrics to record
Define every metric before the test and retain individual cases so that averages do not hide high-impact errors.
| Metric | What to record | Question it answers |
|---|---|---|
| Visual evidence | Expected field, extracted field, and annotated location or reference | Did it read the information that supports the response? |
| Tool decision | Selected tool, omitted call, or unnecessary call | Was the action relevant and authorized? |
| Arguments and execution | Arguments sent, validation, and returned result | Did the tool receive the right data and produce the expected effect? |
| Abstention | Clarification, refusal, or uncertain responses, compared with the evidence | Did it act cautiously when information was missing and proceed when it was available? |
| Latency and cost | Time and available billable units per run and completed task | Is the workflow viable when failures, retries, and review are considered? |
Classify failures before deciding
Incorrectly reading a number or status may result from low resolution, cropping, visual design, or the model’s interpretation. Record the relevant region and how the image was presented; do not automatically attribute every error to a general limitation in vision. If the model answers with a value that does not appear in the image, also classify whether it invented the value, confused it with another field, or inferred it from incomplete context.
Language failures can occur in understanding the request, reading visual content, or producing the final response. Keep these stages separate. The model may understand the request correctly while transcribing a screenshot label incorrectly, or identify the text correctly while misunderstanding which action the user wants. Recording the language of each component helps reveal patterns in mixed-language tasks.
For tool use, distinguish an unnecessary choice, the wrong tool, a malformed schema, valid but incorrect arguments, and a final response that contradicts the returned result. This classification helps identify whether to change tool design, instructions, pre-action checks, or authorization rules. Selecting the right tool does not make up for incorrect arguments.
A convincing result in these tests does not certify universal safety either. The corpus covers only the cases it contains and the configuration that was tested. For a deployment with meaningful consequences, behavioral testing should be accompanied by access limits, argument validation, traceability, and human review appropriate to the risk. These are operational recommendations, not a provider guarantee.
Acceptance criteria and limits of the conclusion
Acceptance criteria should reflect the impact of each task and be set before results are reviewed. For a low-risk operation, an organization might allow a response to be reviewed before a record is changed. For an irreversible action or one with external effects, an unverified call may be unacceptable even if the overall success rate appears high. This protocol does not prescribe a universal threshold: each team must define which errors are tolerable and what controls will contain them.
A prudent decision may assign Command A+ bounded tasks when a representative corpus shows adequate extraction, valid arguments, reasonable abstentions, and a cost compatible with the workflow, always under the planned controls. If errors occur in certain languages, formats, or ambiguous states, those cases can be routed to human review or kept out of scope. The conclusion should name the approved conditions, not claim that the model is generally reliable.
Review Cohere’s model page and guides again when closing the experiment to confirm the current model ID, available channels, applicable formats and limits, and how to configure the request. If a different channel is evaluated or the version changes, the conclusions do not automatically transfer. A published-weights repository may help reproduce local tests under its license conditions, but it does not make local results equivalent to those obtained through an API.
The proposed test also cannot determine how the model will behave with images, languages, tools, or risk levels that were not included. Research on multilingual multimodal reasoning and visual tool chaining can inform corpus design, but it does not replace a current evaluation of the model in the target environment. The strongest conclusion is bounded: what worked, with which configuration, under what conditions, and which errors prevent broader use.
Open questions
- The supplied documentation identifies the model as command-a-plus-05-2026, but the current model ID and access channels should be confirmed when the evaluation is finalized.
- The complete list of 48 languages and public performance results broken down by language for the described tasks are not provided here.
- Image formats, input limits, and parameters compatible with a single request must be checked in the current documentation for the selected channel and configuration.
- The supplied sources do not demonstrate Command A+ performance on tasks combining vision, multiple languages, and tool use.
- The specific cost depends on the channel and current rate; it must be checked for the measured use case rather than inferred from the general billing model.
- Results from published weights and results obtained through an API should not be treated as equivalent without checking configuration, hardware, and conditions.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction