This is not a direct comparison between two models
Claude Sonnet 5 and Cohere Embed 4 do not solve the same task within a retrieval-augmented generation, or RAG, architecture. The former can occupy the generation layer: it receives an instruction and retrieved evidence, then produces an answer. The latter can occupy the representation and retrieval layer: it transforms inputs into vectors to locate related content before the generator answers. Presenting them as though one were an alternative to the other obscures the cause of results and leads to incorrectly attributed architectural changes.
The operational question is not which of the two is “better,” but which combination of retriever and generator provides localized evidence and grounded answers for a particular corpus, set of queries, and constraints. A fluent answer does not prove that the index found the right source. Likewise, having the right passage appear among the results does not prove that the generative model interpreted, cited, and qualified it correctly.
Anthropic’s documentation identifies the API model as `claude-sonnet-5`. Cohere’s documentation, in turn, describes `embed-v4.0` and requires an input type to be declared, including values for search documents, search queries, and images. These facts make it possible to fix the variables of the experiment, but they do not by themselves establish an advantage on a given enterprise corpus.
The evaluation must treat the system as a chain of observable decisions: document conversion, segmentation, indexing, retrieval, context assembly, generation, validation, and presentation. If only the final answer is measured, a team cannot know whether an error comes from a missing page, an overly small chunk, a poorly vectorized query, truncated context, or an unsupported generated claim.
What each layer can contribute, and what its specifications do not prove
Cohere presents Embed Multimodal v4 as a family of unified multimodal embeddings. Its documentation indicates retrieval between text and images and between text and mixed modalities, as well as an advertised context of up to 128k and Matryoshka dimensions of 256, 512, 1024, and 1536. In a RAG system over PDFs containing text, tables, and images, those properties justify testing coherent representations of different evidence types. They do not prove that a specific dimension, or even a multimodal embedding, improves every document collection.
The selected dimension is an experimental variable, not an implementation detail. It affects storage, search latency, and potentially the neighborhood behavior of the index. So does the way a page is represented: as extracted text, as a page image, as a table region, as a combination of those objects, or as several units connected through metadata. A reproducible test must preserve that decision across all branches that are not explicitly evaluating indexing.
Claude Sonnet 5 should be evaluated as a consumer of evidence, not as a silent replacement for retrieval. A correct response when the evidence did not appear in the context may come from the model’s prior knowledge, a fortunate match, or a defect in the pipeline logs. In an application that requires answers grounded only in the corpus, that case does not validate retrieval: it must be labeled as an answer not attributable to the document set, or excluded from the grounded-answer metric.
Generative models can also fail even when they receive the right passages. They may omit a condition from a table, merge two incompatible sources, interpret a number outside its scope, or include a citation that points to a related source but does not support the claim. Citations are therefore evaluable objects, not interface decoration.
Layers, observables, and failures that must not be confused
| Layer | Output that must be retained | Observable failure | Usual decision |
|---|---|---|---|
| Conversion and chunking | Page, region, extracted text, and metadata | Evidence disappears or loses structure | Review parser, OCR, and segmentation |
| Embedding and index | Vector, modality, dimension, and index version | Evidence exists but does not appear in results | Review representation, filter, embedding, or retrieval |
| Generation | Exact context, instruction, and response | Evidence was retrieved but is ignored or distorted | Review model, prompt, and output limit |
| Validation | Claim, citation, and support verdict | The citation does not support what was said | Add verification and human review |
Build a corpus that enables diagnosis, not merely scoring
The evaluation corpus must be frozen before configurations are compared. Include native and scanned PDFs, predominantly textual pages, tables with headers and footnotes, images containing relevant information, and documents that combine those elements. The aim is not to reproduce every variation of enterprise documentation, but to state precisely which variations have been included and to prevent an apparent improvement from depending on documents added or corrected halfway through the test.
Every query needs a reference answer and, above all, a definition of sufficient evidence. That evidence may be a page, a table region, several distributed passages, or a combination of text and image. Record stable identifiers for the document, version, page, region, and chunk. When evidence is distributed, the retrieval judgment must require the necessary set, not merely one of its components.
It is useful to stratify queries by evidentiary pattern. Single-evidence queries reveal whether the retriever finds a localized fact. Multi-source queries reveal whether it delivers the necessary pieces without losing their relationship. Contradictory queries test whether the system preserves provenance and contradiction rather than choosing a source arbitrarily. Queries with no evidence test whether the assistant abstains instead of completing the answer with outside knowledge or an unjustified inference.
Reserve a final split that is not used to tune chunk size, overlap, dimensions, result count, instructions, or validation rules. If a decision is made after observing the test split, that split is no longer an independent measurement. Sample size, distribution by stratum, and the annotation protocol should be published with the results.
Fix the conditions before running the factorial matrix
The comparison requires an explicit baseline pipeline. Keep the parser or OCR, page-resolution policy, chunking, overlap, metadata, vector database, filters, search method, number of retrieved results, reranking—if present—context assembly, prompt, citation policy, maximum output, and retries fixed. If any of these change between branches, the effect cannot be cleanly attributed to Cohere Embed 4 or Claude Sonnet 5.
Also document the access channel, query date and time, region where applicable, SDK or API version, effective model identifier, and submitted parameters. For Cohere Embed 4, record at minimum `embed-v4.0`, the dimension, each object’s modality, and `input_type`. The endpoint documentation distinguishes `search_document`, `search_query`, and `image`; using different types without recording them would change the test. For Claude Sonnet 5, retain the effective request, the unprocessed response, and the context supplied with it.
Define a useful and reproducible baseline. It may be an embedding already adopted by the team and an existing generator, but it must remain unchanged during the experiment. Do not present the baseline as a universal reference: it is only the comparison point for that system. If reranking is introduced, it must remain the same across all four branches or be evaluated as an additional factor, because it could mask or amplify retrieval differences.
The unit of analysis should be the query, but results must also be grouped by document type and evidence pattern. An overall average can hide the fact that a change helps locate visual pages while worsening short textual questions, or that it improves recall while increasing the amount of irrelevant context passed to the generator.
Freezing procedure
- 01Version the corpus, annotations, and query list; assign immutable identifiers.
- 02Run conversion, OCR, and segmentation once for each representation configuration, and save their artifacts.
- 03Declare index, search, context, prompt, and validation parameters before evaluation.
- 04Record every retrieval and every exact context associated with an answer.
- 05Lock adjustments on the final split and analyze that split only once.
The factorial matrix: four configurations, one attribution question
Run four branches on the same queries. The first uses the reference retriever and generator. The second replaces only the embedding layer with Cohere Embed 4 while retaining the baseline generator. The third retains baseline retrieval and uses only Claude Sonnet 5 for generation. The fourth combines Cohere Embed 4 and Claude Sonnet 5. If the system already uses one of them as its reference, redefine the branches so that each factor still has a clearly identified control condition.
This matrix makes it possible to estimate conditional changes, not to establish an absolute hierarchy among providers. The difference between the first and second branches is interpreted as a change associated with retrieval under the baseline generator. The difference between the first and third branches reflects the generation change with evidence retrieved by the baseline. The fourth shows the behavior of the combination, including a possible interaction: the generator may make better use of different evidence, or the retrieval change may not carry through to better answers.
Run several repetitions if generation is not deterministic, and retain seeds or sampling parameters where they are available. For accuracy and abstention tasks, a conservative generation configuration reduces variance and makes auditing easier. Cost comparisons must include amortized indexing, query embedding, retrieval, generation, validation, and retries; reporting generation cost alone would shift spending from one layer to another without representing the real operation.
How to interpret the factorial matrix
| Retrieval | Generation | Analytical use |
|---|---|---|
| Baseline | Baseline | Reference point for the current system |
| Cohere Embed 4 | Baseline | Effect of changing representation and retrieval |
| Baseline | Claude Sonnet 5 | Effect of changing generation over the same evidence |
| Cohere Embed 4 | Claude Sonnet 5 | Result of the combination and interaction between layers |
Measure evidence and answers as distinct outcomes
For retrieval, measure evidence recall: the proportion of queries for which sufficient evidence appears among the results delivered to the generator. Complement it with retrieval precision, or the share of relevant results, because indiscriminately increasing the number of chunks can raise recall and worsen the context. For questions with distributed evidence, measure complete retrieval of the annotated set rather than merely the presence of a partial source.
For the answer, assess factual faithfulness to the documents, coverage of the points required by the query, accuracy of abstention when there is no evidence, and format validity when the application requests a structure such as JSON. Coverage is not the same as faithfulness: an answer may mention every topic and contain a wrong number. Faithfulness is not the same as usefulness either: a strictly supported answer may omit a necessary condition.
Add a supported-citation metric at the level of each verifiable claim. A reviewer, ideally blinded to the configuration, should decide whether the citation or provenance reference supports precisely the associated claim. An automated grader with an audited rubric may be used, but it should be checked against human review on a sample and its disagreements should be retained. RAGBench proposes an explainable evaluation approach that is relevant for decomposing results rather than reducing them to a final score.
Report p50 and p95 latency per answer, not only averages. Separate indexing time from query time, even if the former is included as an amortized cost in an economic metric. The most useful measure for product decisions is often cost and latency per correct grounded answer, provided that the definition of “correct” is published and the denominator includes appropriate abstentions.
Diagnose error patterns
When annotated evidence does not appear among the results, the incident initially belongs to retrieval or to an earlier conversion and indexing stage. Check whether the text or image was available after parsing, whether the indexed unit retained the table context, whether metadata filtered out the document, and whether the query used the declared modality and input type. Changing the generator does not fix the absence of evidence from the context.
When sufficient evidence does appear and the answer is wrong, the incident concerns evidence use, unless context assembly truncated the relevant passage before it was sent. Compare the supplied text with the final claim: this reveals omissions, source blending, excessive inferences, and ignored contradictions. Here, it may be reasonable to investigate the instruction, context structure, output limit, or generative model.
A decorative citation is a separate pattern: the answer includes a real reference, but it does not support the specific proposition. It is especially common when the document contains related terms or when a table is cited at page level while the number comes from another row. Labeling it simply as an incorrect answer makes it impossible to know whether the problem lies in citation generation, provenance granularity, or downstream validation.
Unjustified abstention requires the opposite diagnosis. If sufficient evidence was retrieved and the system said it could not answer, the instruction for using context, the order of results, or the readability of extracted content may have failed. If no evidence existed, abstention is correct even if it is less satisfying to the user. A reliable system must distinguish the two cases and explain its limits without inventing an answer.
Per-query diagnostic tree
- 01Does sufficient evidence exist in the frozen corpus? If not, assess abstention rather than retrieval.
- 02Did the evidence reach the retrieved results? If not, investigate conversion, segmentation, embedding, index, and filters.
- 03Did the evidence reach the generator’s effective context? If not, investigate top-k, reranking, and truncation.
- 04Does the answer match the evidence, and does every citation support it? If not, investigate generation, context composition, and validation.
- 05Was the answer correct without evidence in context? Label it as not attributable to grounded RAG.
Make decisions without extrapolating beyond the experiment
Change the retrieval layer when the main deficit is the absence of relevant evidence, particularly in subsets where text, images, or mixed pages must be found in relation to one another. Before concluding that the embedding is the cause, rule out OCR, chunking, and metadata errors. A table separated from its headers, for example, may fail with any embedding because the indexed unit lost the necessary semantics.
Change or adjust the generative layer when evidence arrives sufficiently and consistently, but the answer uses it poorly. Even then, do not assume that the model name explains the entire change: test citation instructions, context formats, and abstention rules on the development split, then confirm them on the final split. Claude Sonnet 5 should be compared through equivalent requests and complete logs, not selected examples.
Consider reranking if initial recall is reasonable but top results contain too much noise or relegate decisive evidence. Consider human review for high-impact answers, document contradictions, complex tables, or cases where automated validation cannot reach sufficient agreement with reviewers. None of these measures removes the need to measure each layer: it adds another variable that must also be audited.
Publish the experiment card with the results: corpus definition and version, annotation scheme, query distribution, effective configurations, retrieval artifacts, grader criteria, agreement among evaluators, statistical intervals or uncertainty, latency, cost, and representative errors. Do not publish only an aggregate rate or demonstration examples. The retrieval-quality evaluation framework and work on multimodal RAG reinforce the need to observe access to evidence and the answer separately.
Results do not automatically generalize to every enterprise repository. The prevalence of scanned PDFs, languages, OCR quality, table density, permissions, document-update patterns, and cost tolerance can change the decision. The defensible conclusion is local: under documented conditions, one configuration retrieved certain evidence better or produced more faithful answers. That precision is more useful than assigning general capabilities to models that perform different functions.
Open questions
- Provider documentation describes capabilities and identifiers, but does not guarantee relative results on a specific enterprise corpus.
- Behavior may vary with OCR, parser, language, chunking strategy, filters, dimensions, top-k, reranking, and context composition.
- Assessing faithfulness and citation support involves human judgment; the rubric, reviewer agreement, and disagreements should be published.
- No experimental results are provided for the four configurations; this article defines an evaluation protocol, not a performance ranking.
- Availability, pricing, access channels, and effective parameters must be verified and recorded at the time of each run.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction