Cohere's 111B enterprise model - 256K context, 8K output, citations and grounding as API parameters, and non-commercial weights that run on two GPUs.
Why this oneWhat it is actually for
You reach for Command A when the answer has to say where it came from. Grounding is not a prompt pattern here, it is two request parameters: you pass a documents array and the response comes back with citation spans attached, controlled by citation_options. That plus safety_mode and strict_tools makes it the model to pick when a compliance reviewer will read the output. The second reason is deployment: at 111B parameters it fits on two A100s or H100s, Cohere sells private deployments through SageMaker, Azure AI Foundry and Oracle OCI, and the weights are on Hugging Face for non-commercial use - so you can evaluate locally before signing anything. Note it is a March 2025 model, not a 2026 one.
What it isArchitecture, lineage, training
Command A is a dense 111 billion parameter autoregressive transformer. Cohere Labs' model card describes an interleaved attention pattern: three consecutive layers use sliding window attention with a window of 4096 tokens and RoPE position encoding, and every fourth layer uses global attention with no positional embeddings, letting tokens interact across the whole sequence. That is what makes 256K context affordable on two GPUs - most layers only ever look 4096 tokens back, and the periodic global layer carries the long-range signal. It is not a mixture-of-experts model; all 111B parameters are active on every token.
Training is described at a high level only. Cohere states the model is trained for 23 languages and aligned with supervised fine-tuning followed by preference training. The technical report is arXiv 2504.00698. Cohere does not publish the training token count, the data mix or the compute.
Weights are published by Cohere Labs as CohereLabs/c4ai-command-a-03-2025 under CC-BY-NC with the Cohere Lab Acceptable Use Policy - non-commercial only, and the card notes the Hugging Face release is configured for 128K context rather than the 256K the API serves. Commercial use means the API or a paid private deployment.
Lineage matters here because Cohere has moved on. Command A's successor command-a-plus-05-2026 is documented as Cohere's first mixture-of-experts model, folding vision input, agentic behaviour, reasoning and translation into one set of weights, at 128K context and 64K output - and Cohere's pricing page lists it under an Apache 2.0 licence with a free model download. Command A remains Live in Cohere's model table with a longer 256K context, but it is no longer the front of the line.
At a glanceSee it
Grounding documents and interleaved local and global attention feed one 8K-capped, citation-bearing answer.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model string | command-a-03-2025, status Live. Note that Cohere's own Command A docs page contradicts this: its Model ID card renders command-a-plus-05-2026, its successor's string. The models overview table is the one to trust. | Cohere models overview table; Cohere Command A docs page |
| Context window | 256,000 tokens on the API. The Hugging Face release is configured for 128K. | Cohere models table and Command A docs Specifications block; Cohere Labs model card |
| Max output tokens | 8,000 | Cohere models overview table; Command A docs Specifications block |
| Modalities in / out | Text in, text out. The models table lists Modality "Text". | Cohere models overview table |
| Parameters | 111 billion, dense. Hugging Face reports 111,057,580,032 BF16 as a single total, and Cohere describes Command A+ as its "first Mixture of Experts model", which places Command A outside that family. | Cohere Command A docs page; Cohere Labs model card; Hugging Face safetensors index |
| Knowledge cutoff | June 1, 2024 | Cohere Command A docs page, Specifications block |
| Languages | 23, named explicitly in the docs | Cohere Command A docs page |
| Hardware floor | Two GPUs, A100 or H100 | Cohere Command A docs page |
| Weights available | Yes, CohereLabs/c4ai-command-a-03-2025 under CC-BY-NC-4.0 plus the Cohere Lab Acceptable Use Policy. Non-commercial only. The repo is gated - you must be signed in and accept the terms before any file will download, config.json included, so the architecture cannot be checked anonymously. | Cohere Labs Hugging Face model card; Hugging Face repo metadata |
| Endpoint | Chat, v2 at https://api.cohere.com/v2/chat | Cohere v2 Chat API reference |
| Release | March 2025, from the 03-2025 version stamp | Cohere models table; model card version |
| Training token count and data mix | Not disclosed | No Cohere page states either |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
max_tokens | Output cap. | Optional. If unset, defaults to the model's maximum output token limit - 8,000 for Command A. Cohere does not document what happens to a request that asks for more. | Set too low and finish_reason comes back MAX_TOKENS with a truncated answer. |
temperature | Randomness. | Non-negative float. Default 0.3. | Cohere ships cold. Raising it alone does less than you expect because p also defaults below 1. |
p | Nucleus sampling. Note the name is p, not top_p. | 0.01 to 0.99. Default 0.75. | The default already discards the tail. Push to 0.95 before blaming temperature for flat output. |
k | Top-k. Name is k, not top_k. | 0 to 500. Default 0, which disables it. | If both k and p are set, p is applied after k. |
seed | Best-effort determinism. | Integer, 0 to 18446744073709552000 | Cohere states determinism cannot be totally guaranteed. Useful for regression tests, not for audit. |
frequency_penalty | Penalises tokens in proportion to how often they have appeared. | 0.0 to 1.0. Default 0.0. | Range tops out at 1.0, not 2.0. Small moves matter. |
presence_penalty | Flat penalty on any token already seen. | 0.0 to 1.0. Default 0.0. | Use for topic drift, not repetition; frequency penalty is the repetition tool. |
stop_sequences | Halt strings. | Up to 5 strings | The stop sequence itself is excluded from the returned text. |
tools | Function definitions the model may call. | Array of ToolV2 objects | Cohere notes Command A is deliberately good at not calling tools it does not need. |
tool_choice | Forces or forbids tool use. | REQUIRED or NONE. Omit to let the model decide. Command R7B and newer only. | There is no "call this specific tool" option - only all-or-nothing. |
strict_tools | Beta. Forces tool calls to follow the schema exactly. | Boolean | The first few requests with a new tool set are slower while the schema is compiled. |
documents | Grounding sources for RAG. | List of strings, or objects with content and metadata | This is what turns on citation spans. Also blocks response_format. |
citation_options | Controls citation generation. | Object. Modes include ENABLED, DISABLED, FAST, ACCURATE, OFF. Enabled by default. | FAST trades citation precision for latency; OFF removes the spans entirely. |
response_format | Forces JSON output, optionally against a JSON Schema. | {"type": "json_object"}, with optional json_schema | Not supported in combination with documents or tools. Without a schema, nesting is capped at 5 levels. Your prompt must explicitly say to generate JSON or the model can emit an unbounded character stream. |
safety_mode | Selects the injected safety instruction. | CONTEXTUAL default, STRICT, or OFF | Not configurable in combination with tools or documents. |
thinking | Reasoning configuration on the v2 Chat schema. | Object with type enabled or disabled and token_budget | Cohere's reasoning documentation demonstrates this only against command-a-reasoning-08-2025. Its effect on command-a-03-2025 is not documented - treat it as not applicable here. |
In practice three things decide how Command A behaves. documents plus citation_options is the whole reason to use this model, and it is also the constraint that rules out response_format - you get grounded prose with citations or you get JSON, not both, so structured RAG has to go through tools with strict_tools instead. p at 0.75 is the second: it is the lowest default nucleus value among mainstream APIs and it is why Command A reads more clipped than its temperature suggests. Third is the system prompt - Cohere warns the model is verbose and markdown-happy by default and tells you to instruct it otherwise.
SamplingShaping the output distribution
Command A ships conservative on two axes at once. Temperature defaults to 0.3 and p defaults to 0.75, so before you touch anything the distribution has already had its tail cut at 75% of probability mass and then been sharpened. That compounding is the thing people miss: raising temperature to 0.9 while leaving p at 0.75 gets you much less variety than the same move on an API whose nucleus default is 1.0. If you want genuinely diverse output, raise p toward 0.95 first and temperature second. k is off by default at 0, and Cohere documents the interaction order plainly - when both are enabled, p acts after k. A sensible starting point for grounded question answering is to leave both defaults alone; they are tuned for exactly that. For drafting or ideation, p 0.95 with temperature 0.7 is the first thing to try. For extraction, temperature 0 with the default p, and add seed if you need repeatable test fixtures - Cohere says determinism is best-effort, not guaranteed.
ReasoningThinking, effort and budgets
Command A has no reasoning mode. Cohere's reasoning models are hybrid and gated behind a separate model string: command-a-reasoning-08-2025, which shares the 256K context but raises max output to 32K to make room for thinking. On that model, thinking is on by default and you disable it by sending thinking with type set to disabled. The budget is a token count via token_budget, and Cohere's own guidance is to leave it unlimited; if you must cap it, leave at least 1K tokens for the answer and use roughly 31K as the budget against the 32K output ceiling. When the budget is exhausted the model jumps straight to the final answer. Thinking content is visible - it arrives as content blocks with type of thinking, separate from the text blocks. On Command A itself, your options are to prompt for stepwise working in the visible answer, or switch model strings.
ToolsFunction calling and server tools
Tool support is strong and the constraints are specific. tools takes function definitions, tool_choice takes only REQUIRED or NONE - there is no way to name one tool and force it - and strict_tools is a beta flag that makes the model obey your schema exactly, at the cost of slower first requests with a new tool set. Cohere highlights that Command A is trained to avoid calling tools it does not need, which matters in agent loops where over-eager calling is the usual failure. Structured output has a real trap: response_format is documented as not supported in combination with either documents or tools. So the two things you most want together - grounded answers and machine-readable output - cannot both come from response_format. Use strict_tools with a schema-shaped tool instead. Two more: when you do use response_format, your prompt must explicitly instruct the model to generate JSON or it can loop until it runs out of context; and safety_mode is silently unavailable alongside tools or documents.
CostPrice, caching, batching, what drives the bill
Command A's list price is on its documentation page, not on Cohere's pricing page. As of 2026-07-25 the Generative models section of cohere.com/pricing lists only Command A+, Command R, Command R7B and Transcribe - Command A is absent, while its docs page still shows a Pricing block. Take that as a signal about where Cohere is steering new usage. There is no published prompt-caching discount and no published batch discount for the Chat endpoint, though the v2 response does report usage.cached_tokens, "the number of prompt tokens that hit the inference cache" - so a cache exists, it just has no published rate attached. On reading your bill: Cohere's How Does Cohere's Pricing Work? page - not the tokens-and-tokenizers guide - explains that Cohere adds tokens under the hood and that "since these are tokens you don't have control over, you are not charged for them." Reconcile against usage.billed_units, not usage.tokens; the two differ in the same response.
| Model | Input per 1M | Output per 1M | Where listed |
|---|---|---|---|
| Command A | $2.5 | $10 | Cohere Command A docs page, Pricing block |
| Command A+ | $0 | $0 | cohere.com/pricing, listed as "API key / Model download" with an Apache 2.0 licence |
| Command R | $0.15 | $0.60 | cohere.com/pricing |
| Command R7B | $0.0375 | $0.15 | cohere.com/pricing |
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Cohere API | Yes | command-a-03-2025 on Chat v2, Chat v1 and a Chat Completions compatibility endpoint. Base URL https://api.cohere.com/v2/chat. |
| Amazon Bedrock | No | Cohere's platform table lists the Bedrock model ID for command-a-03-2025 as "(Coming Soon)". No Cohere SKU appears in the AWS Price List for us-east-1 pulled 2026-07-23. |
| Amazon SageMaker | Yes | Listed as "Unique per deployment" - the model ID is created when you deploy. |
| Microsoft Foundry / Azure AI Foundry | Yes | Listed as "Unique per deployment". |
| Oracle OCI Generative AI | Yes | Model ID cohere.command-a-03-2025. |
| Google Vertex AI | Unverified | Not listed in Cohere's platform table as of 2026-07-25. |
| Hugging Face weights | Yes | CohereLabs/c4ai-command-a-03-2025, CC-BY-NC. Non-commercial use only, and configured for 128K rather than 256K. |
| Self-hosting | Yes | Two A100s or H100s for the open weights under the non-commercial licence; commercial self-hosting goes through a Cohere private deployment. |
StrengthsWhat it is good at
- Grounding and citation are API parameters, not prompt craft:
documentspluscitation_optionswith ENABLED, FAST, ACCURATE and OFF modes. - 256K context served on two A100 or H100 GPUs, achieved by using sliding-window attention with a 4096 window on three of every four layers and global attention on the fourth.
- The full sampling and penalty surface is present and documented with exact ranges -
seed,k,p, both penalties,logprobs- which is rare among enterprise-focused APIs. - Weights are downloadable for evaluation under CC-BY-NC, so you can test locally before committing to a private deployment.
safety_modegives an explicit CONTEXTUAL, STRICT or OFF switch rather than an opaque built-in filter.
LimitsWhere it falls down
response_formatis not supported alongsidedocumentsortools, so grounded structured output has to be routed throughstrict_toolsinstead.- Knowledge cutoff is June 1, 2024 - old enough that anything time-sensitive must come in through
documents. - Text only, 8K output cap, no reasoning mode. Each of those is a separate Cohere model: Command A Vision, Command A Reasoning, Command A+.
- The weights licence is CC-BY-NC. Downloading and self-hosting for a commercial product is not permitted; that path requires a paid private deployment.
- Cohere's own Command A documentation page is internally inconsistent as of 2026-07-25 - its Model ID widget shows
command-a-plus-05-2026and its capability chips include Reasoning and Image Inputs, both of which contradict the models table entry for Command A. Copy the model string from the models table, not the product page.
Against its neighboursHow it compares
The most direct comparison is Cohere's own Command A+ (command-a-plus-05-2026): mixture-of-experts, vision input, reasoning and translation in one model, 128K context and 64K output, listed on Cohere's pricing page under Apache 2.0 with a free model download. It beats Command A on almost everything except raw context length, where Command A's 256K is double. If you are starting fresh, start there and fall back to Command A only if you need the longer window. Against Amazon Nova Pro, the split is operational: Nova Pro is cheaper per token and native to AWS, Command A gives you citations, a grounding parameter, safety modes and portable weights. Against Mistral Large 3, both target regulated European and enterprise deployment with private hosting; the differentiator to test is citation quality on your own corpus, since that is where Cohere has spent its effort.
Getting startedThe smallest call that works
curl https://api.cohere.com/v2/chat \
-H "Authorization: Bearer $CO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "command-a-03-2025",
"messages": [
{"role": "user", "content": "How long is parental leave?"}
],
"documents": [
"Policy 4.2 - parental leave is 16 weeks at full pay."
],
"max_tokens": 512
}'Add a system message first - Cohere warns the model is verbose and uses markdown by default, and a one-line instruction to answer plainly changes the output more than any sampling change will. After that, raise p from its 0.75 default if answers feel clipped, and check the response for citation spans before you build any UI around them.
SourcesWhere every claim above came from
- An Overview of Cohere's Models - the models table giving
command-a-03-2025, Live status, 256k context, 8k max output, Text modality, plus the per-platform table with Bedrock "(Coming Soon)", SageMaker, Azure AI Foundry and Oracle OCI entries, and the Command A+ row at 128k/64k. - Command A - Cohere docs - 111B parameters, 256,000 context, 8,000 max output, June 1 2024 knowledge cutoff, two-GPU hardware floor, $2.5 and $10 per 1M pricing, 23 languages, and the note that the model is chatty by default.
- Cohere Chat v2 API reference - every request parameter with exact name, default and range: temperature 0.3, p 0.75 with 0.01-0.99 bounds, k 0-500, both penalties 0.0-1.0, seed bounds, stop_sequences limit of 5, tool_choice values, strict_tools, citation_options, response_format restrictions, safety_mode, thinking, priority, and the endpoint URL.
- Reasoning Capabilities - Cohere docs - thinking enabled by default on reasoning models, token_budget guidance, the 31K budget recommendation, and that the example runs against command-a-reasoning-08-2025.
- CohereLabs/c4ai-command-a-03-2025 model card - CC-BY-NC licence, the three-sliding-window-plus-one-global attention pattern with a 4096 window, the 128K Hugging Face configuration, SFT plus preference training, 23 languages, arXiv 2504.00698.
- Cohere Pricing - read as raw HTML because the tables render client-side. Generative models listed are Command A+ at zero cost with an Apache 2.0 licence and model download, Command R at $0.15/$0.60, Command R7B at $0.0375/$0.15, plus Transcribe. Command A is not listed.
- Could not confirm: Command A's training token count, data mix or compute; any prompt-caching or batch discount; Vertex AI availability. Cohere's Command A docs page currently displays
command-a-plus-05-2026in its Model ID widget and lists Reasoning and Image Inputs among its capabilities, contradicting the models table - this page follows the models table.