Ilustración editorial para Command A+ vs. Command R 08-2024: cómo decidir una migración para un asistente RAG multilingüe
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The question is not which model appears more capable, but what changes in the complete product

A migration from Command R 08-2024 to Command A+ should be evaluated as a system change, not as an isolated replacement of a model identifier. In a retrieval-augmented generation enterprise assistant, the outcome depends on the query, indexing, retriever, available documents, instructions, output validation, retries, and human intervention. If any of those elements changes between tests, it will not be possible to attribute a difference to the model.

Cohere’s documentation places both models in a context relevant to this use case: Command R 08-2024 is positioned for retrieval workflows, citations, tools, and multilingual use; Command A+ documents citations, tools, and structured outputs, in addition to an input modality that includes text and images. That modality difference is an available capability, but it does not demonstrate an advantage in a test where all queries, evidence, and responses are text-only.

The useful decision is therefore specific: under the assistant’s real contract, does Command A+ verifiably increase the share of usable responses that are faithful to the evidence and operationally compatible? The evaluation must also ask whether that improvement offsets any integration, latency, or effective-cost change. Without a shared measurement, provider-announced capabilities are context for designing the test, not evidence of improvement for a particular application.

02

Published capabilities to normalize before measuring

According to Cohere’s model inventory and model pages, Command R 08-2024 uses the `command-r-08-2024` identifier, accepts text, has a 128,000-token context window, and a 4,000-token maximum output. Command A+ is identified as `command-a-plus-05-2026`, accepts text and image input, generates text, has a 128,000-token context window, and a 64,000-token maximum output. Status and availability should be recorded again on the evaluation cutoff date because these are attributes that may change in the provider’s documentation.

These differences require two questions to remain separate. The first is the common comparison: both models receive the same textual input, the same retrieved context, and the same practical response limit. The second is an evaluation of expanded capabilities, such as images or long outputs. Combining them would favor the model with a broader contract even if that breadth is not enabled in production.

The configuration that can alter generation behavior must also be fixed. If the deployment uses reasoning, tool calls, citations, or a JSON format, the configuration must be equivalent within the space both models share. Do not assume that settings with the same name produce the same behavior. Record the request, raw response, client version, and effective parameters so discrepancies can be investigated.

Variables that must remain equal and variables requiring a separate test

ElementTreatment in the common text comparisonTreatment in an additional evaluation
Query, corpus, and retrieverIdentical for both modelsIdentical unless the question studies multimodal retrieval
Output lengthOne common practical limit, below the shared maximumA specific test if the product needs lengthy responses
Image inputExcludedA separate Command A+ pilot with its own data, security, and evaluation
ToolsThe same simulated catalog and approval policyAn additional test only if the tool contract changes
API and messagesIdeally shared; otherwise, record the differenceValidate API migration as a separate hypothesis
03

Build a corpus that represents costly errors, not only easy questions

The evaluation set should be derived from real tasks, anonymized when necessary, and maintain strict separation between development and final testing. For every query, retain the language, intent, domain, length, expected answer when one exists, admissible evidence passages, and expected or prohibited action. Retrievable documents must be versioned: a silent index update invalidates the comparison.

A multilingual sample should include the languages the service actually supports, rather than machine translations of a single question set. It should incorporate short and long queries, local terminology, ambiguities, and requests that mix languages. Quality must be segmented by language rather than reduced to a global average: an aggregate improvement can conceal a meaningful regression for a particular region or team.

The corpus needs deliberate negative cases. Include insufficient evidence, conflicting documents, outdated data labeled as such, tables converted to text, out-of-scope requests, and action requests that must wait for human approval. Correct abstention is a useful result; a convincing answer without documentary support is not. Simulated actions make it possible to evaluate tool selection without executing real changes in customer, billing, or knowledge systems.

04

Design a common, auditable harness

The harness should run every case against both models using the same set of already retrieved documents or, if the goal is to measure end-to-end retrieval, the same retriever, indexes, filters, search query, and number of passages. These are two different measurements. Supplying identical passages to the model isolates generation and citation attribution; running the retriever measures the complete product behavior, but introduces additional sources of variation.

Use a semantically equivalent template that establishes response language, source priority, the obligation to abstain, citation format, field schema, and the rule not to execute actions. Control the retry budget: for example, a retry following invalid JSON must be applied under exactly the same conditions. Counting an answer as successful only after repairing or retrying it, without reflecting that in cost and latency, distorts the conclusion.

Tools should be deterministic and simulated. Each can return predefined results, log arguments, and block side effects. The decision to score is whether the authorized tool was selected, whether arguments are valid, and whether the assistant left the action pending review. Do not infer safety from a high correct-selection rate: permission and approval policy remain the application’s responsibility.

Reproducible batch execution

  1. 01Freeze the corpus, index or supplied passages, simulated tools, and client configuration.
  2. 02Assign an immutable identifier to every case and run both models in randomized order to reduce temporal effects.
  3. 03Store the request, raw response, citations, tool calls, errors, retries, token usage, and timestamps.
  4. 04Automatically validate the output contract before sending the result to blinded human evaluation.
  5. 05Evaluate faithfulness, accuracy, and abstention without revealing which model produced the response.
  6. 06Segment results, calculate acceptance outcomes, and manually review the highest-impact regressions.
05

Measure the usable end-to-end outcome

Evidence retrieval should be evaluated before assessing prose. For every important claim, determine whether sufficient evidence existed among the supplied passages and whether the response selected relevant passages. Citation faithfulness is more demanding: every citation must support the specific claim it accompanies, rather than merely address the same topic. When the corpus contains contradictions, the evaluation should check whether the assistant expresses uncertainty or correctly applies the defined priority rule.

Structured accuracy must be measured field by field. An answer can be valid JSON and still contain an incorrect identifier, date, amount, or status. Distinguish syntactic validity, schema compliance, completeness, semantic accuracy, and compatibility with downstream logic. If a field is not supported by the context, the expected behavior may be null, an uncertainty marker, or an abstention, according to the contract fixed in advance.

For abstentions, measure precision and coverage. Penalize both invention when evidence is absent and unjustified refusal when evidence is sufficient. For tools, measure selection, arguments, and respect for human approval. Finally, connect technical performance to operations: calculate p50, p95, and p99 latency per usable response, and effective cost including tokens, failed calls, retries, and responses that do not pass validation. A model that is faster per request can be less efficient if it requires more repair.

Decision matrix for metrics

MetricUnit of analysisSuggested acceptance criterionRisk if omitted
Evidence faithfulnessClaim and citationEach important claim is supported by the cited passagePlausible but unverifiable answers
Structured accuracyFieldCorrect value that is valid under the schemaSilent failures in automations
AbstentionCase with and without evidenceAnswers when appropriate and abstains when support is absentHallucinations or excessive refusal
ToolsSimulated requestCorrect tool and arguments; does not execute without approvalIncorrect or unauthorized actions
OperationsAccepted responseLatency and cost measured after validation and retriesOptimization based on unusable requests
06

Treat integration compatibility as a testable hypothesis

It is not advisable to promise that changing the model without modifying the client will work merely because both models belong to the same provider. The migration guide between API V1 and V2 documents differences in messages, response fields, streaming, documents, citations, and tool calls, as well as V1 features that V2 does not support. If the migration includes an API change, the source of a regression may be the integration contract rather than the model.

Structured outputs require particular caution. Cohere’s documentation includes Command A+ and Command R 08-2024 among models compatible with Structured Outputs and describes the use of JSON Schema and strict tools. It also states that Structured Outputs JSON is not supported in RAG mode. Consequently, an assistant that needs both RAG citations and JSON must not assume the two functions can be combined in one request mode; it must turn that combination into an explicit integration test.

Before a pilot, create contract tests for ordinary requests, streaming, limit errors, incomplete responses, empty citations, tools without valid arguments, and cancellations. Compare the objects consumed by the application, not only the visible text. A small incompatibility, such as a missing optional field or a different tool-call representation, can stop a downstream flow even when the linguistic response is correct.

07

Interpret results by segment, not only by average

Present results for the full set and for segments defined before opening the data: language, context length, retrieval difficulty, document conflict, tabular extraction, and action proposal. Report the number of cases per segment, invalid responses, abstentions, and excluded cases along with their reasons. Excluding only one model’s failures or changing the reference after seeing responses biases the comparison.

An accuracy improvement may not offset a regression in abstention, citations, or cost. For example, Command A+ may generate longer answers under a broad limit, but that difference should not count as an improvement if the product requires a concise answer and the excess degrades the interface or review process. Likewise, the fact that Command A+ accepts images permits no conclusion about a text-only corpus.

Human review should be blinded to the model and based on a scoring guide with edge examples. When two reviewers disagree, a third can resolve the case or mark it as ambiguous. The disagreement rate should be retained as a signal of metric uncertainty. In high-impact matters, review by domain experts is preferable to automated evaluation based solely on textual similarity.

Reading a regression

  1. 01Check that the case used the same context, instructions, output limit, and harness version.
  2. 02Distinguish between evidence not retrieved, evidence retrieved but not used, an unfaithful citation, incorrect extraction, and a formatting failure.
  3. 03Reproduce the case without retries and then with the production policy to quantify the operational effect.
  4. 04Review whether the failure is concentrated in a language, document type, or integration path.
  5. 05Turn the confirmed cause into a regression test before changing configuration or deploying.
08

Decide whether to retain, adopt, or pilot

Keeping Command R 08-2024 is a reasonable decision if it meets established quality and operational thresholds and Command A+ does not deliver a material, consistent, attributable improvement. It may also be the prudent choice if the new contract requires API or validation changes whose risk has not yet been tested. The relative age of a release is not, by itself, a replacement criterion.

Adopting Command A+ is justified when it exceeds predefined thresholds across the complete outcome: reliable evidence and citations, correct fields, appropriate abstention, safe tool use, verified compatibility, and acceptable cost or latency. The improvement must survive priority segments and an integration regression test. It is advisable to define in advance which degradation, if any, is unacceptable; for example, a drop in citation faithfulness for a workflow that must be auditable.

A limited pilot is appropriate when there are positive indications but uncertainty remains around real traffic, concurrent loads, minority languages, or integration. Route a controlled fraction of eligible requests, retain human review, and enable immediate rollback. Do not use the pilot to discover an undefined contract: its success criteria, duration, population, and stopping conditions should be established before users are exposed.

Practical decision rule

Observed outcomeIndicative decisionAdditional condition
Consistent improvement in validated quality and no critical operational regressionGradually adopt Command A+Pass contract tests and have rollback available
Advantage limited to some segments or production uncertaintyLimited pilotInstrumentation, human review, and stopping criteria
No verifiable improvement or regression in critical criteriaKeep Command R 08-2024Record findings and repeat only after a relevant change
Need for images or a different output contractSeparate evaluationDo not extrapolate from the text protocol
09

Validate again when images or other contract changes are enabled

Command A+ image input opens a possibility that Command R 08-2024, according to the supplied documentation, does not share. That capability requires a new evaluation, not an automatic extension of text-only results. The corpus should include representative images, transcriptions or ground truth, criteria for unreadable data, and controls for privacy, retention, and access. The team must also decide what counts as citable evidence when part of the information comes from an image.

Likewise, a larger maximum output is valuable only if a product need requires longer responses or transformations. Test the use case with real limits, truncation mechanisms, budget, and quality evaluation. Do not turn a published maximum capability into a configuration recommendation without observing its effect in the system.

This protocol does not predict a winner. The available sources are provider documentation and describe published capabilities, limits, and contracts; they do not provide independent results for an organization’s corpus. The responsible conclusion should be stated only after running the harness, retaining the artifacts, and declaring the segments that did not have enough size or review to support a decision.

Open questions

  • The supplied documentation comes from the provider and describes published capabilities, not independent quality, latency, cost, or reliability results on a specific corpus.
  • Availability status, identifiers, limits, and API contracts can change and should be recorded again on the evaluation cutoff date.
  • No execution data, pricing, region, load, supported languages, or human-review results were provided, so it is not possible to declare a winning model.
  • The exact combination of retrieval, citations, and JSON output depends on the selected mode and integration and must be verified through contract tests.
10

Keep exploring

10

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction