Ilustración editorial para Command R 08-2024 y sus modos de seguridad: qué control ofrecen y qué debes probar
Imagen generada con gpt-image-2.5-sunburst para InferamaSource ↗
01

What safety_mode controls—and what it does not

Command R 08-2024 is a Cohere model version identified by the model name command-r-08-2024. When you call Cohere’s chat endpoints, the safety_mode parameter lets you choose between modes that change the safety instructions included in the request. In other words, it is a configuration for how the model is guided during an interaction, not an independent certification of the model or a guarantee that an application is safe.

That distinction matters in practice. An acceptable result on a simple prompt describes only that combination of model, endpoint, parameters, and instructions. On its own, it does not show that an application will withstand adversarial inputs, handle untrusted documents correctly, or block every harmful output. Those properties require testing the complete system and applying product controls appropriate to the risk.

The provider’s available documentation should not be read as evidence that a mode eliminates all problematic content. The model card provides context about safety evaluations and limitations, but it does not demonstrate how effective each mode is for command-r-08-2024. It is therefore better to describe the modes as a configuration option to evaluate, not as a complete defense.

02

STRICT, CONTEXTUAL, and NONE: differences and endpoint names

The Chat V1 reference documents the values STRICT, CONTEXTUAL, and NONE, and marks safety_mode as beta. The current Chat V2 reference lists STRICT, CONTEXTUAL, and OFF. NONE and OFF therefore appear as names in different references; do not assume they are interchangeable values within the same call. Before configuring an integration, check the specification for the endpoint you use.

Cohere historically presented the modes as different levels of safety guidance. STRICT prioritizes stricter application of those instructions; CONTEXTUAL applies them with regard to the request’s context; and NONE is the option without the set of instructions added by the mode. This does not mean that NONE turns the model into a system with no protections at all: the provider’s announcement says that some protections cannot be disabled.

The purpose of CONTEXTUAL is to avoid reducing safety to a mechanical reaction to isolated words. However, the name’s implication of contextual interpretation is not proof that the model will always correctly distinguish a legitimate request from a harmful one. Teams should test this using examples from their own domain, including cases where the right interpretation is genuinely uncertain.

The beta label in V1 is relevant to change management: it is a reason to review the documentation and rerun tests before updating an integration. The V2 reference has its own set of values, but the sources provided do not adequately clarify whether V1’s beta status carries over, or whether availability details are identical across deployment channels. Do not automatically transfer the status of one endpoint to another.

An operational reading of the modes

A summary for designing tests. It is not a guarantee of the outcome of any particular call.

ModeName in the referenceOperational interpretationWhat to test
STRICTV1 and V2Strict application of safety instructions.Whether it correctly refuses harmful requests and whether it blocks legitimate sensitive tasks.
CONTEXTUALV1 and V2Safety instructions applied with regard to context.Whether it distinguishes permitted uses from harmful requests in closely related cases.
NONE / OFFNONE in V1; OFF in V2The set of instructions associated with the mode is not added; this does not imply that all protections are absent.What behavior the model retains and how the application responds without relying on this mode.
03

The tools and documents limitation changes what a test can establish

The Chat V1 reference documents that safety_mode cannot be combined with tools, tool_results, and documents. The Chat V2 reference also indicates a limitation on combining the setting with tools and documents. For a team using retrieval-augmented generation (RAG) or tool calls, this is an operational constraint: do not present an isolated test of the mode as direct evidence of how the complete workflow will behave.

The documented incompatibility also does not mean that tools or documents are unsafe, or that the model cannot use them in any configuration. It says that the listed parameters cannot be combined in a call as described by those references. The integration must follow the endpoint’s supported request format, and the components and their interactions should be validated separately.

A useful evaluation separates at least two questions. First: how does the model respond with safety_mode in a compatible, controlled call? Second: what happens in the real application when documents are retrieved, tool results are sent, or the application’s own instructions are added? The answer to the first question does not replace the answer to the second. A retrieved text may also contain instructions, incorrect information, or harmful material; the application should treat it as content to evaluate, not as a safety policy.

The sources provided describe parameter restrictions in the Chat V1 and Chat V2 references. They do not detail every available combination in every deployment channel, nor do they support a claim that every Cohere-compatible service offers identical capabilities. Confirm the specification for your actual environment before designing the test.

04

A reproducible evaluation protocol

A comparative evaluation should define in advance what outcome counts as correct for each case. If the team changes its criteria after seeing the responses, it may end up favoring the mode that confirms its prior expectations. Maintain a labeled set of cases with a rationale for each label, and run each case using the same model version, endpoint, and parameters, except for the variable being compared.

Include permitted but sensitive requests: for example, preventive security guidance, an educational question about a high-risk topic, or a request for support that calls for a careful tone. Add clearly harmful requests according to the product’s policy, as well as ambiguous cases where the model should ask for context, limit its answer, or refuse part of the request. These are examples for building an evaluation set; the sources do not publish a test suite specifically designed to measure the modes in Command R 08-2024.

Test variations in language and phrasing: paraphrases, spelling errors, indirect expressions, and changes of language that matter to your audience. Simply translating one case literally is not enough, because the intended meaning may not carry over between languages. If the product receives content in multiple languages, evaluate each language with competent speakers or with criteria reviewed by specialists.

Run each case more than once. Responses can vary between executions, and a single sample cannot establish consistency. Keep the exact inputs, expected outcomes, full responses, evaluation decisions, and call context. When the model, API, or application instructions change, rerun the set and compare versions.

Steps for comparing modes

Keep all variables other than safety_mode constant. If the endpoint does not permit a required combination, record that test as a separate evaluation of the production configuration.

  1. 01Define a response policy and acceptance criteria before running tests.
  2. 02Prepare permitted but sensitive, harmful, ambiguous, and language-variation cases; document why each belongs in its category.
  3. 03Record the endpoint, model identifier, mode, parameters, application instructions, and execution date.
  4. 04Repeat each case and retain inputs and outputs so discrepancies can be reviewed.
  5. 05Evaluate results using consistent criteria, and keep simple calls separate from workflows involving documents or tools.
  6. 06Review errors, update the test set, and validate again before using the results to make a product decision.
05

Metrics that help interpret the results

Count correct refusals: cases that should have been refused and received a refusal or a safe response under the defined criteria. Also count false refusals: permitted requests that were blocked or made unusable. A high refusal rate does not, by itself, demonstrate better safety if it is achieved by rejecting many legitimate tasks.

Also record unsafe responses: cases that should have been refused or limited but in which the system provided disallowed content. Analyze product-instruction compliance, consistency across executions, and differences between modes separately. If a response is partly correct, use intermediate categories defined in advance rather than forcing it into a simple “pass” or “fail.”

Break down results by request type, language, mode, and configuration. An overall average may conceal different behavior on permitted sensitive questions and ambiguous requests. For workflows involving documents or tools, also record what content was retrieved, which tool was called, and what information was returned to the model, while following your organization’s privacy and retention requirements.

Metrics are evidence for a decision, not a universal certification. A small sample, an overly predictable test set, or imprecise labeling rules can create a false impression of reliability. Report the test set’s size and composition, the number of repetitions, and disagreements between evaluators. When results are ambiguous or have significant consequences for people, include specialist human review.

Metrics and the decisions they inform

Interpret each metric alongside the case types and acceptance criteria; avoid choosing a mode based on a single number.

MeasureQuestion it answersSignal to investigate
Correct refusalsDoes the system refuse or limit cases the policy identifies as disallowed?Harmful cases answered without the expected limitation.
False refusalsDoes the system block legitimate, sensitive requests?Permitted tasks that cannot be completed usefully.
Instruction complianceDoes it follow the response rules defined for the application?Instructions omitted, contradicted, or applied inconsistently.
ConsistencyDoes it produce comparable results across repeated executions?Meaningful changes between responses to the same input.
Configuration differenceDoes the result vary between a simple call and an integrated workflow?Differences that might be attributable to documents, tools, or additional instructions.
06

Record the environment and keep external controls

For results to be interpretable, record the exact model identifier—including command-r-08-2024 where applicable—the endpoint, the safety_mode value, and the date. Also record other generation parameters, system and application instructions, the input language, and whether documents, tools, or tool results were included. The date and integration version help identify cases where two apparently comparable runs did not use the same environment.

The provider documentation does not answer every deployment question raised by this guide. In particular, the available evidence is insufficient to confirm whether the beta status of V1 does or does not apply to V2, whether the parameter-combination limitation has identical scope across all channels, or whether a protocol can be reproduced unchanged in every compatible environment. Verify those points against the current reference for your endpoint and your actual configuration before interpreting results.

Keep your own controls even if a test favors one mode. These may include defining permitted uses, access controls for tools, data validation, action limits, monitoring, incident-response procedures, and human review for high-impact decisions. Which controls are appropriate depends on the product and its risks; a sufficient configuration cannot be inferred from the mode’s name alone.

As a practical rule, choose the mode that meets your application’s policy with an acceptable balance between unsafe responses and false refusals, based on representative tests. If the result depends on using documents or tools, evaluate the production workflow separately and confirm which parameters it supports. Repeat the evaluation when the model, endpoint, instructions, or architecture changes. This ties the decision about Command R 08-2024 to evidence from the system actually in use, rather than relying on a mode label.

Open questions

  • The sources provided do not establish whether the beta status of safety_mode in Chat V1 carries over to Chat V2.
  • The cited documentation does not establish that the same parameter combinations are available across all deployment channels.
  • No published evaluation is provided that isolates the effectiveness of STRICT, CONTEXTUAL, and NONE/OFF specifically for Command R 08-2024.
  • The practical equivalence of NONE in V1 and OFF in V2 should not be assumed beyond the difference in naming described in the references.
07

Keep exploring

07

Sources consulted

03

Corrections and transparency

If you spot incorrect or outdated information, send us a correction with the page and source we should review.

Submit a correction