Primary use stated or inferred cautiously from official documentation.
Command A+
Enterprise RAG, multilingual use, and private deployment.
The essential
Maximum input capacity when published by the source.
Documented maximum generation limit.
Verified availability channels.
Primary functional family and verified specialties.
Declared terms for API access, model weights, or self-hosting.
Limits and integration
| API ID | command-a-plus-05-2026 |
|---|---|
| Model type | Text and reasoning |
| Access model | Paid proprietary |
| License | Proprietary |
| Deployment | Hosted API |
| Release | 20 MAY 2026 |
| Knowledge cutoff | 1 APR 2025 |
| Entrance | Text · Image |
| Exit | Text |
| Context window | 128,000 tokens |
| maximum output | 64.000 tokens |
| Reasoning | Supported |
| Published tools | Features · Citations · RAG |
| Structured Outings | Supported |
| Batch processing | Not published |
| Prompt cache | Not published |
| Fine-tuning | Private deployment and open weights |
| Verified platforms | Cohere API · Azure AI Foundry |
Enterprise RAG, multilingual use, and private deployment.
No public per-token rate is published for private production.
What it can do
Orientation
Enterprise RAG, multilingual use, and private deployment.
Context
128.000 tokens · 64.000 tokens
Tools and integration
Features · Citations · RAG
Access
Cohere API · Model Vault · private deployment
Documented cost
| Concept | Worth | Unit/condition |
|---|---|---|
| Standard input | Free up to the quota limit | Standard API rate |
| Cached input | Not published | Reading reused prefixes |
| Cache write or storage | Not published | The condition varies by provider |
| Standard output | Free up to the quota limit | May include reasoning tokens |
| Batch input | Not published | Asynchronous processing |
| Batch output | Not published | Asynchronous processing |
Production use is offered through Cohere Model Vault and enterprise agreements.
i Prices change and may depend on level, region or context length. Check the source before making a decision.
How to read the results
A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.
| Benchmark | Result | Metric | Source |
|---|---|---|---|
| Throughput frente a Command A Reasoning | +110 % | Throughput | View source ↗ |
| Latencia frente a Command A Reasoning | −30 % | Latency | View source ↗ |
Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.
Chronology
Command A+
Version added to Inferama's verified catalog.
Related analysis
Command A+: How to Test Whether Vision, Multilingualism, and Tool Use Work Together
A list of capabilities does not show that they work well in combination. This protocol proposes testing Command A+ with images, requests in multiple languages, and simulated tools, while recording successes, errors, latency, and cost.
25 Sep 2026 ↗ COMPARATIVACommand A+ vs. Command R 08-2024: How to Decide on a Migration for a Multilingual RAG Assistant
Changing models does not necessarily improve an enterprise assistant. This comparison proposes a reproducible protocol to determine whether Command A+ delivers a net improvement over Command R 08-2024 when both operate on the same corpus, retriever, output contract, and human review process.
22 Sep 2026 ↗