Ilustración editorial para Cohere Embed 4: cómo migrar un índice multimodal sin mezclar espacios vectoriales
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

The unit of analysis is the model together with its index

Changing an embedding model is not the same as swapping out a single interchangeable component in a search system. The model transforms documents and queries into vectors; the index organizes those vectors and makes it possible to compare them. Retrieval depends on both sides sharing a compatible representation and on processing the query with an appropriate configuration. When evaluating Cohere Embed 4—identified as `embed-v4.0` in the provider’s documentation—the unit of analysis should therefore be the combination of model, input configuration, vectors, index, and comparison rule.

The operational consequence matters: vectors created by an earlier model should not be added to the same search space as vectors generated by Embed 4 without validation. The fact that both models produce vectors of the same length does not show that their coordinates have the same meaning. Nor is it enough to change the output dimension in the storage system. If you replace the model, the prudent approach is to regenerate representations for both documents and queries with a compatible configuration, while keeping the old index available during evaluation.

This guide sets out a method for deciding whether to migrate, add a multimodal path, or keep the current solution. It does not assume that Embed 4 will outperform the existing system: specifications describe declared capabilities, and usage documentation explains how to configure the model, but the effect on relevance depends on the corpus, queries, and implementation. The focus is retrieval—not comparing generative models or attributing quality to the final answer from a RAG system.

02

What Cohere states about Embed 4

Cohere presents Embed 4 as a multimodal model that can generate representations from text, images, and mixed inputs. Documented mixed-input examples include PDF pages containing both text and images. This makes it possible to consider use cases that are not limited to searching text extracted from files: a page could contribute both textual and visual information to its representation. However, support for a modality does not guarantee that every document in that modality will be interpreted correctly or that retrieval will improve for a particular task.

Cohere’s published information states that the model offers output dimensions of 256, 512, 1024, and 1536, as well as a 128k context. These are provider specifications, not a promise that every access channel accepts identical formats, limits, or parameters. Cohere’s API, Amazon Bedrock, and Oracle Cloud Infrastructure, for example, are different interfaces. Before designing a migration, check the documentation for the service you will actually use and record the exact model name, region or service where applicable, current limits, and request format.

Dimension is part of the contract between the embedding generator and the index: it determines the size of each vector and therefore affects compatibility, storage, and search operations. A smaller dimension may reduce the amount of data associated with vectors, but you should not infer that it will preserve retrieval quality. Likewise, a larger dimension does not guarantee better performance on your corpus. Compare alternatives using the same queries, relevance judgments, and operating conditions.

Embed 4 has an `input_type` parameter associated with the purpose of the input. In Cohere’s embedding documentation, `search_document` identifies document inputs and `search_query` identifies search queries. In a retrieval workflow, using the appropriate type for each side is part of the configuration that must remain consistent across indexing and querying. Do not treat it as an optional descriptive label: follow the values supported by the selected endpoint and check how they apply to text, images, or mixed inputs.

Configuration decisions to validate

Specifications help formulate tests; they do not replace measurement on your team’s corpus.

DecisionWhat to verifyWhat the specification alone cannot establish
ModalityWhich formats the access channel accepts and how text, images, or mixed inputs are sent.That every visual file or PDF will produce correct retrieval.
DimensionWhich values the integration allows and which are compatible with the test index.That a larger dimension will necessarily improve relevance.
`input_type`Which value is appropriate for documents and queries on the specific endpoint.That vectors produced with different configurations are interchangeable.
Context and limitsCurrent length, size, and format limits for the selected service.That a limit documented for one service is identical on another.
03

Why you should not mix vectors from different models

A vector is a numerical representation produced by a model under a particular configuration. Vector search usually ranks candidates using a similarity or distance measure. For that comparison to make sense, document and query vectors must be generated compatibly with the method used to retrieve them. Identical dimensions show only that the numeric lists have the same length; they do not establish that corresponding positions can be compared across two models.

The distance metric is also part of the index configuration. Do not change the comparison metric during migration without checking how it affects results. Cohere’s documentation describes similarity measures for embeddings, but the specific choice depends on the integration and index. The team should record which measure the current system uses, which the candidate will use, and whether the search engine normalizes vectors or applies additional transformations.

Keep the separation in both storage and evaluation. Label each vector with its model, dimension, processed modality, and configuration version so that you can reconstruct how it was produced. If the old and new indexes run in parallel, each query should generate the vector corresponding to each index. Do not send a query embedded by one model to an index built with another just to simplify routing, unless a deliberate test has shown that the combination is valid for the use case.

A common mistake is to evaluate only the answer generated by a RAG application. A convincing answer can conceal the fact that relevant documents never appeared among the top results; it may also rely on the generative model’s prior knowledge. To evaluate the embedding layer, inspect what the search system retrieved and compare it with relevance judgments that are independent of the text later produced by the generative model.

04

Migrate documents and queries with traceability

A controlled migration begins with an inventory of the current index. Record its model and identifier, dimension, distance method, chunking strategy, prior transformations, input type, covered languages, and retained metadata. For multimodal use, also note which modalities are extracted or sent, how pages are identified, and how each vector relates to the original chunk. Without this inventory, a difference in results could be caused by the model, chunking, or an unnoticed preprocessing change.

Next, build the candidate index as a new version. Regenerate document embeddings, retain stable keys that link each vector to its source, and record any item that could not be processed. Do not overwrite the old representations during the test. For each run, store the model, dimension, input type, date, and process version. This makes it possible to reproduce a comparison and detect whether two apparently identical tests used different configurations.

The query should follow an equivalent path. The search service identifies which index is being queried and uses the intended input type for queries. During a period of parallel operation, you can send the same logical query to both the old and candidate paths, but each must generate its own embedding. If a later stage combines results, measure its effects separately; otherwise, it will be difficult to attribute an improvement or regression to Embed 4.

For images and PDF pages, document how files are handled by the specific access channel. Cohere’s documentation shows a semantic-search workflow with PDF pages and mixed content, but exact formats and limits must be checked in the selected API or platform. Do not automatically apply Cohere’s rules to Bedrock or OCI, or vice versa. Nor should you interpret the stated context as permission to send any file without regard to size, structure, or format restrictions.

Safe migration sequence

Keeping the current version available makes comparison and rollback possible without mixing representation spaces.

  1. 01Inventory the current index’s model, dimension, metric, chunking, preprocessing, and metadata.
  2. 02Freeze a reproducible copy of the corpus and select documents representing relevant languages and modalities.
  3. 03Create a separate candidate index and recalculate document vectors using the configuration verified for Embed 4.
  4. 04Generate query vectors with the corresponding query configuration and run them against the candidate index.
  5. 05Compare results with the current index, recording processing failures, latency, and operational consumption.
  6. 06Promote the candidate only if it meets agreed criteria; retain the previous index and a rollback path.
05

Test protocol: corpus, queries, and relevance

Before comparing models, fix an evaluation corpus and avoid changing it between runs. Include common, difficult, and infrequently queried documents; if search covers more than one language, include examples in each. For multimodal evaluation, adding a few images is not enough: identify tasks where visual information is necessary, documents containing both text and images, pages with tables, and cases where OCR-extracted content may be incomplete. The aim is to represent intended use, not to build a sample that favors a model in advance.

Prepare real queries or queries derived from observable needs, and record which documents or chunks should count as relevant. Judgments can be binary or graded, but apply the same rules when comparing the current system and the candidate. Include direct and ambiguous queries, rare terms, proper names, and questions whose answers depend on an image or table. A query should not become a relevance judgment simply because the current system answers it correctly: define what you expect to retrieve.

The analysis should examine both positions and retrieved sets. Recall@k can show whether relevant items appear among the first k results; precision@k can show what proportion of the initial results are judged relevant; and nDCG@k can be useful when judgments distinguish degrees of relevance. These are possible metrics, not a guarantee that any single indicator will summarize system utility. Agree in advance on which values of k and which promotion criteria matter to the real user experience.

Break down results by language, modality, and query type. An overall improvement can hide a decline in a less common language or in documents containing images. Likewise, an acceptable average does not eliminate critical failure cases. Keep examples of queries for which the new model retrieves something different, review them with domain specialists, and distinguish indexing errors, chunking errors, configuration errors, and relevance issues in the embedding itself.

Minimum evaluation matrix

Complete the matrix with corpus data and criteria agreed by the team; it does not assume results for Embed 4.

SegmentExamples to includeSignals to review
LanguageEach language with a significant presence in actual use.Relevance among top results and types of query that fail.
TextDirect and ambiguous queries, plus queries with rare terms.Recall@k, precision@k, or a graded measure defined by the team.
ImageSearches where a visual feature is necessary.Whether relevant pages are retrieved and whether visual information adds value.
Mixed PDFPages with text, images, tables, or complex layouts.Processing errors, loss of context, and retrieval of the correct chunk.
OperationsRepresentative queries and workloads from production.Latency, storage volume, consumption, and service failures.
06

Operational metrics and the effects of dimension

Retrieval quality is not the only decision criterion. Measure latency under comparable conditions, both for generating embeddings and for search, and record the resources needed to process the corpus and maintain the index. Separate the cost of creating or rebuilding embeddings from the cost of serving routine queries. Prices and timings depend on the provider, access channel, input size, selected dimension, and workload; they cannot be derived from the model card alone.

Compare available dimensions in a test that keeps other factors constant. Check the effective storage footprint in the index, the implications for transfer, and retrieval behavior. If the provider describes a Matryoshka strategy or the option to choose among several dimensions, treat it as a configuration choice to measure, not as automatic proof of equivalence across sizes. Reducing dimension may ease some operational loads, but it could change which documents rank first.

Also record the proportion of inputs rejected, truncated, or processed differently than expected. In a multimodal test, the percentage of documents that never made it into the index can alter retrieval metrics and give a misleading picture of the model. Report processing coverage separately from retrieval quality on valid cases, as well as overall results that account for failures. A fair comparison should not silently exclude difficult documents for one configuration.

07

Promotion, rollback, and documentation

Define before running the test what results would be sufficient to promote the candidate index. Criteria might combine retrieval thresholds by segment, no regressions on critical queries, acceptable latency and cost limits, and a minimum level of processing coverage. There is no universal threshold: it depends on the importance of each use case, the service level, and the cost of errors. The decision should rest on observable evidence, not an impression that demo answers seem better.

Rollback should be part of the migration design. Retain the previous index, its configuration, and the mapping between document identifiers and vectors. If promotion fails, the query route should return to the previous version without adding candidate vectors to the old index or losing traceability. During a gradual transition, identify each result with the index and configuration that produced it; avoid combining outputs from both versions without a tested merging policy.

Document the result, including limitations: languages with small samples, unevaluated modalities, inputs the access channel does not accept, and differences between providers. If the team operates directly through Cohere, Amazon Bedrock, or OCI, record which documentation and limits apply to that integration. A model card for a model offered on one platform should not become a general claim about every endpoint using the Embed 4 name.

A migration is defensible if someone else can reconstruct which corpus was tested, with what model and parameters, which queries and judgments were used, which metrics were calculated, and how the decision was made. Cohere’s public documentation is useful for identifying declared capabilities and parameters; validating retrieval performance in your own environment remains the team’s responsibility.

Evaluation exit criteria

The final decision should consider retrieval quality, operations, and the ability to roll back.

  1. 01Approve the exact endpoint configuration, dimension, input types, and distance metric.
  2. 02Review metrics and examples by language, modality, and query class—not just the overall average.
  3. 03Confirm that costs, latency, and processing coverage remain within agreed limits.
  4. 04Verify that the rollback path preserves the current index and its traceability.
  5. 05Publish the decision with results, test limitations, and conditions for repeating the evaluation.
08

What can be concluded, and what requires evaluation of your own

Cohere’s documentation states that Embed 4 supports text, images, and mixed inputs, offers selectable dimensions, and has a large context; its examples include search over PDF pages. The provider’s API and guides describe parameters relevant to retrieval, such as `input_type`. This information helps plan a test and identify which configurations need to be checked. It does not show that an index built with another model can be reused directly, that a particular dimension is optimal, or that retrieval will improve on a specific document set.

The useful question is not whether Embed 4 is better in the abstract, but whether a defined `embed-v4.0` configuration produces acceptable retrieval results for the team’s documents, languages, modalities, and operational constraints. Answering that requires building a separate candidate index, recalculating documents and queries compatibly, evaluating against relevance judgments, and recording cost and latency. Only after that comparison does it make sense to decide whether to replace the current system or add a multimodal path.

For people exploring models and tools in Inferama’s discovery area, this approach provides a basis for comparison that goes beyond a model specification sheet. The Cohere Embed 4 page and information about Cohere as an organization can help locate the model and its documentation; the operational choice, however, should be justified by reproducible results from the index and real queries.

Open questions

  • Specifications and limits may differ between Cohere’s API and third-party integrations; check the current documentation for the selected endpoint.
  • Provider sources describe capabilities but do not establish an independent retrieval improvement for a particular corpus.
  • The effect of each dimension on quality, storage, and latency must be measured under the team’s actual workload.
  • Performance for each language, modality, and document type cannot be inferred from an aggregate metric or from the stated ability to process multimodal inputs.
09

Keep exploring

09

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction