Home › Frontier Models › Mistral Large 3
Model · Reference

Mistral Large 3

EU data residency, open option

In one line

A 675B-total, 41B-active mixture-of-experts with a 256k window, Apache 2.0 weights, and API pricing of $0.50 in and $1.50 out per million tokens.

Why this oneWhat it is actually for

You reach for Mistral Large 3 when you want frontier-shaped behaviour without being locked to one vendor's endpoint. It is one of very few models at this scale released under Apache 2.0 — not a community licence with a user-count clause, but the real thing — so the same weights you rent from the API can be pulled onto your own B200 or H200 node when a customer, a regulator or a procurement team says the data cannot leave. The price makes that credible rather than theoretical: $0.50 per million in and $1.50 out is well under most frontier rates. Choose it for long-document work, enterprise agents and tool use. Do not choose it for hard chain-of-thought maths — Mistral says so itself.

What it isArchitecture, lineage, training

Mistral Large 3 is a granular sparse mixture-of-experts, released 2 December 2025, trained from scratch on 3,000 H200s. The headline split is 41B active parameters out of 675B total, which the model card decomposes further: a 673B-parameter language model with 39B active, plus a 2.5B vision encoder.

The architecture numbers come from params.json in mistralai/Mistral-Large-3-675B-Instruct-2512 — Mistral ships its own format, so there is no config.json in that repository. The language model has n_layers 61 at dim 7168, with n_heads 128 and head_dim 192. Attention is the latent variety rather than plain GQA: q_lora_rank 1536, kv_lora_rank 512, qk_nope_head_dim 128, qk_rope_head_dim 64, v_head_dim 128, so the KV state is compressed to a 512-wide latent instead of stored per head. The MoE block gives num_experts 128 with num_experts_per_tok 4 and num_shared_experts 1, at expert_hidden_dim 4096; first_k_dense_replace 3 means the first three layers stay dense before routing begins. Four of 128 routed experts is what "granular" means here — a far sparser draw than the eight-of-eight-times-seven era.

Positional handling is rope_theta 10000 over an original_max_position_embeddings of 8192, extended by YaRN with factor 36 to a max_position_embeddings of 294912 — the 256k window Mistral advertises sits inside that table with headroom. Vocabulary is 131072 on the Tekken tokenizer. The vision encoder is 48 layers at hidden size 1664, patch size 14, images up to 1540 pixels on the long edge, processed by a Pixtral-family processor. The default instruct repository ships FP8 block-quantised weights; NVFP4 and BF16 variants are published separately.

At a glanceSee it

Mistral Large 3 diagram

Four of 128 experts fire per token, so only 41B of 675B parameters are billed as work.

CapacityContext, output and what fits

FactValueSource
API model idmistral-large-2512, aliased as mistral-large-latestdocs.mistral.ai model card and mistral.ai/pricing/api
Context window256k tokens. params.json declares max_position_embeddings 294912 via YaRN factor 36 — Mistral supports 256kModel card, docs.mistral.ai and Hugging Face; params.json
Max output tokensNot disclosed as a separate cap. The API states prompt plus max_tokens cannot exceed the context lengthdocs.mistral.ai API reference, chat completions
Modalities inText and imagesHugging Face model card, Key Features
Modalities outTextHugging Face model card
Knowledge cutoffNot published as a spec. The SYSTEM_PROMPT.txt shipped in the repository states the knowledge base was last updated 2023-10-01, which is difficult to reconcile with a December 2025 release — treat it as unconfirmedSYSTEM_PROMPT.txt in the Hugging Face repository
Weights availableYes — Base, Instruct in FP8, plus NVFP4, BF16 and Eagle variantsmistralai collection on Hugging Face
LicenceApache 2.0Repository metadata and mistral.ai/news/mistral-3
Release date2 December 2025docs.mistral.ai model card and the Mistral 3 announcement
Total / active parameters675B total, 41B active. Language model 673B / 39B, vision encoder 2.5BHugging Face model card
Experts128 routed, 4 per token, 1 shared; first 3 layers denseparams.json, moe block
Vocabulary131,072, Tekken tokenizerparams.json and tekken.json
LanguagesRepository metadata lists en, fr, es, de, it, pt, nl, zh, ja, ko, ar; the announcement claims 40+ native languagesHugging Face repository metadata and mistral.ai/news/mistral-3

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
temperatureScales logits before samplingDefault varies by model and is not published. Mistral advises below 0.1 for productionAbove about 0.3 tool-argument construction and long-document extraction start drifting. This is the knob Mistral names explicitly
top_pNucleus cut-off, an alternative to temperatureNumber or nullDocumented as an alternative rather than a companion to temperature; changing both at once makes results hard to attribute
max_tokensCeiling on generated tokensInteger or nullPrompt plus max_tokens must fit the 256k window. There is no separate output cap to work around
stopHalts generation on a stringString, array of strings, or nullThe stop text is not returned. Useful for delimiter-framed output that you then parse
random_seedRequests deterministic samplingInteger or nullSame seed plus same request should reproduce; it does not make an MoE router bit-identical across backends
response_formatConstrains output shape{"type": "text"} by default; also json_object and json_schemajson_schema is the one to use — json_object guarantees valid JSON but not your fields
toolsFunction definitions available to the modelArray or nullMistral's card advises keeping the set minimal; accuracy degrades when the model is handed a large catalogue
tool_choiceForces or forbids tool useauto, none, any, required, or a named toolany and required guarantee a call happens, which is how you stop the model answering from memory instead of calling
parallel_tool_callsAllows several tool calls in one turnBoolean, default trueSet false when your tools share mutable state and must run in order
presence_penaltyPenalises tokens already used, onceNumber or nullEncourages vocabulary diversity; on JSON output it starts corrupting repeated key names
frequency_penaltyPenalises tokens by usage countNumber or nullSuppresses looping in long generations; too high and technical terms get avoided mid-document
prompt_cache_keyMarks a prefix for cache reuseString or null; cached tokens billed at 10%The single largest cost lever on repetitive workloads — a 90% cut on the input side
safe_promptInjects Mistral's safety preambleBoolean, default falseTurning it on prepends a system instruction you did not write, which can conflict with your own system prompt
nCompletions per requestNot supported on mistral-large-2512Mistral's sampling guide states this model does not support N completions. Issue separate requests, or use a smaller model for candidate generation
prediction / Predicted OutputsSupplies an expected completion to speed up editingSupported. Prediction object or nullMistral's model card lists Predicted Outputs among this model's capabilities on /v1/chat/completions. Worth using when you are rewriting a large document or patching code and most of the output is already known — it cuts latency without changing what the model is allowed to produce
prompt_mode and reasoning_effortEnable native reasoning tracesEndpoint accepts prompt_mode: "reasoning" and effort levels none through xhighUnverified for this model. Mistral's reasoning docs name only mistral-small-latest and mistral-medium-3-5, and the Large 3 card calls it "not a dedicated reasoning model"

Three things matter here. Temperature is the one Mistral itself calls out — below 0.1 for daily-driver and production use, which is a stricter recommendation than most vendors publish and worth following literally. prompt_cache_key is the cost lever: cached input bills at 10%, so a stable system prompt and tool catalogue pays for itself immediately. And tool_choice plus a deliberately small tools array is how you get reliable agentic behaviour, because the card is explicit that overloading the model with tools degrades it. Note also that top_k is not a parameter of this API at all — it only exists if you self-host.

SamplingShaping the output distribution

Mistral gives an unusually concrete recommendation for a model this size: use a temperature below 0.1 for daily-driver and production environments, and explore higher values only for creative work. That is stricter than the 0.7-ish defaults people carry over from other APIs, and it reflects what the model is built for — long-document comprehension, agentic tool use and enterprise workflows, where a confident single answer beats a varied one. The API documents top_p as an alternative to temperature rather than a companion, and publishes no default value for either, so a request that omits both is running on an unstated per-model default. A sensible starting point is temperature 0.05 with top_p left alone, moving to 0.7 and above only for drafting and ideation. Two mechanical constraints shape this further: n is not supported on mistral-large-2512, so you cannot generate and rank candidates in one call, and there is no top_k, so nucleus sampling is the only pool control the hosted API exposes. Self-hosting on vLLM restores both.

ReasoningThinking, effort and budgets

No — and Mistral says so plainly rather than leaving you to infer it. The Hugging Face card's known-issues section opens with "Not a dedicated reasoning model", noting that dedicated reasoning models outperform it in strict reasoning use cases, and the launch post positions it at number two in the open-source non-reasoning category on LMArena. The chat completions endpoint does accept prompt_mode set to "reasoning" and a reasoning_effort ladder running none, minimal, low, medium, high, xhigh, and when effort is high the response content becomes a list of chunks with a thinking type alongside a text type. But Mistral's native-reasoning documentation names only mistral-small-latest and mistral-medium-3-5; whether those fields do anything on mistral-large-2512 could not be confirmed from a primary source as of 2026-07-25. What you do instead: reach for Magistral or a Ministral 3 reasoning variant when the task is competition maths or multi-step proof, and use Large 3 for everything around it.

ToolsFunction calling and server tools

Tool use is the strong suit — the card calls out native function calling and JSON output as best-in-class, and the model card's feature grid confirms which endpoints actually light up for this model: Chat Completions, Function Calling, Agents and Conversations, Built-In Tools, Structured Outputs, Prefix, Document QnA and Batching. It also confirms what does not: Predicted Outputs, FIM, OCR, embeddings, moderation, transcription and text-to-speech are all disabled for Large 3, whatever the platform offers elsewhere. parallel_tool_calls defaults to true, and tool_choice accepts auto, none, any, required or a named function. Structured output should go through response_format with json_schema rather than json_object. The failure mode Mistral names explicitly is tool sprawl: keep the set well-defined and limited to the minimum the use case needs. The second one is temperature — a request left at 0.7 will produce plausible-looking but wrong tool arguments far more often than one at 0.05.

CostPrice, caching, batching, what drives the bill

List price from Mistral's own API pricing page, read 2026-07-25:

ItemPrice
Input, per 1M tokens$0.50
Output, per 1M tokens$1.50
Cached input90% discount — cached tokens billed at 10% of the input rate
Batch processing50% discount

Two structural facts drive the actual bill. First, the discounts stack against different workloads: prompt_cache_key cuts the input side by 90% on repetitive prefixes, and the batch endpoint halves everything for anything that tolerates latency. Second, there are no reasoning tokens to pay for, because this is not a reasoning model — the output count is the answer, not the answer plus a hidden trace. That makes cost forecasting much more honest here than on effort-ladder models. The tokenizer is Tekken with a 131,072-token vocabulary, which is dense on European languages and code but offers no unusual advantage on English prose. Self-hosting has no per-token price at all: FP8 fits a single node of B200s or H200s, NVFP4 a node of H100s or A100s.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Mistral API / La PlateformeYesPOST https://api.mistral.ai/v1/chat/completions, model id mistral-large-2512
Amazon BedrockYesMistral's cloud-deployment doc lists Mistral Large 3 (25.12); access via the Bedrock Converse API
Microsoft Foundry / Azure AIYesListed on Mistral's Azure deployment page as Mistral Large 3 (25.12), serverless or real-time endpoints
Google Vertex AINoMistral's Vertex page lists only Mistral Medium 3, Codestral 2, Mistral OCR and Mistral Small — Large 3 is absent
Hugging Face weightsYesBase, Instruct FP8, NVFP4, BF16 and Eagle variants under mistralai, Apache 2.0
Self-hostingYesvLLM >= 1.12.0 with mistral_common >= 1.8.6. Mistral states the model was not added to transformers
Hugging Face InferenceUnverifiedHugging Face is named as a distribution channel in the launch post; hosted inference for a 682GB checkpoint was not confirmed

StrengthsWhat it is good at

  • Apache 2.0 at 675B parameters — a genuinely permissive licence at frontier scale, with the base checkpoint published alongside the instruct one for custom post-training.
  • $0.50 in and $1.50 out per million, with a 90% cached-input discount and a 50% batch discount, puts it well below most models it is benchmarked against.
  • The same weights run on the vendor API, Bedrock, Azure AI and your own hardware, so a data-residency requirement changes your deployment without changing your model.
  • Deployment envelope is published rather than guessed: FP8 on one node of B200s or H200s, NVFP4 on one node of H100s or A100s.
  • Mistral publishes real architecture in params.json — 128 experts, 4 active, latent attention ranks, YaRN factor — which almost no frontier-scale vendor does.

LimitsWhere it falls down

  • Not a reasoning model, by Mistral's own admission: the card states dedicated reasoning models outperform it on strict reasoning tasks, and the launch post ranks it in the non-reasoning category.
  • Behind vision-first models on multimodal tasks, again stated on the card — the 2.5B encoder is a capable add-on, not the point of the model.
  • n completions are not supported on mistral-large-2512, and Predicted Outputs are disabled for it, so two common latency and sampling tricks are unavailable.
  • Not available on Google Vertex AI, unlike Mistral Medium 3 and Codestral 2.
  • Mistral did not add the model to transformers, so self-hosting means vLLM 1.12.0 or newer, and 682GB of shards to move before you serve anything.

Against its neighboursHow it compares

Against Qwen3 Max and DeepSeek V4 Pro, the case for Mistral Large 3 is licence and jurisdiction: Apache 2.0 from a French vendor, with Azure and Bedrock endpoints inside EU regions, which is often the whole reason it wins a procurement rather than a benchmark. On raw reasoning it does not claim to lead — Mistral positions it in the non-reasoning tier deliberately. Against GPT-5.6 Sol the comparison is a different shape entirely: Sol is rented, Large 3 can be owned, and the fallback path when a customer refuses a US API matters more than a few benchmark points. The nearest thing to a like-for-like is the Llama 5 row elsewhere on this site — except that model does not exist, which leaves Mistral Large 3 as the strongest genuinely open-weight frontier model you can download today.

Getting startedThe smallest call that works

code
curl https://api.mistral.ai/v1/chat/completions \
  -H "Authorization: Bearer $MISTRAL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistral-large-2512",
    "messages": [
      {"role": "user", "content": "Summarise the termination clause in this contract."}
    ],
    "temperature": 0.05
  }'

The temperature above is already the change most people need to make — Mistral recommends below 0.1 for production, and the default is unpublished. Add prompt_cache_key next, as soon as your system prompt stops changing, for a 90% cut on cached input. Then response_format with a json_schema if you are parsing. Do not reach for n; this model rejects it.

SourcesWhere every claim above came from

  • Mistral Large 3 — Mistral Docs model card — release date 2 December 2025, model id mistral-large-2512, 256k context, $0.5 / $1.5 per M, and the feature grid showing which endpoints are enabled and disabled for this model
  • mistralai/Mistral-Large-3-675B-Instruct-2512 model card — 675B / 41B split, 673B+39B language model plus 2.5B vision encoder, Apache 2.0, temperature below 0.1 recommendation, vLLM 1.12.0 floor, known issues
  • Mistral Large 3 params.json — dim 7168, 61 layers, 128 heads, head_dim 192, kv_lora_rank 512, q_lora_rank 1536, 128 experts with 4 per token plus 1 shared, first_k_dense_replace 3, vocab 131072, rope_theta 10000, YaRN factor 36, max_position_embeddings 294912, FP8 block quantisation, 48-layer vision encoder
  • Mistral Large 3 SYSTEM_PROMPT.txt — the shipped system prompt, including its knowledge-base date claim
  • Introducing Mistral 3 — Mistral AI — 2 December 2025 launch, Apache 2.0, 3,000 H200s, LMArena placement in the non-reasoning category, list of launch availability partners
  • Mistral API pricing — $0.5 input and $1.5 output per million for Mistral Large 3, batch at 50% off, cached input at 90% off
  • Mistral API reference — chat completions — full request parameter list with types, defaults and allowed values including tool_choice, response_format, prompt_cache_key, safe_prompt, prompt_mode and reasoning_effort
  • Mistral Docs — Sampling — states that mistral-large-2512 does not support N completions
  • Mistral Docs — Reasoning — names mistral-small-latest and mistral-medium-3-5 as the reasoning models; Mistral Large 3 is not listed
  • Mistral Docs — Amazon Bedrock and Azure AI and Vertex AI — per-cloud model availability lists used for the surfaces table
  • Could not confirm: a separate maximum output token limit for this model; a published default value for temperature or top_p; a knowledge cutoff, the only figure available being the 2023-10-01 date inside SYSTEM_PROMPT.txt, which is inconsistent with a December 2025 release; and whether prompt_mode or reasoning_effort have any effect on mistral-large-2512. The site's row also listed pricing as unset and the release year as 2026 — the vendor's own pages give 2 December 2025 and $0.50 / $1.50, and I have used those. Benchmark charts on the model card are images and were not read.
A living map of modern AI — kept current every morning