Home › Frontier Models › Kimi K3
Model · Reference

Kimi K3

Cost-sensitive coding/agents

In one line

Moonshot's 1M-context flagship: always thinking, and the only model on this board that rejects temperature and top_p outright.

Why this oneWhat it is actually for

You reach for Kimi K3 when the job is long-horizon coding or agent work, the input is enormous, and you want frontier-class output at a fifth of what the US flagships charge for output tokens. The 1M window means a repository or a full document corpus goes in without chunking, and the cache-hit rate of $0.30 per million makes a fixed system prompt or codebase preamble nearly free to re-send. You reach for something else if you need deterministic sampling — K3 does not accept temperature or top_p at all — if you need a batch discount, or if your compliance posture cannot route data to a Chinese API endpoint. There is no self-host escape hatch yet either.

What it isArchitecture, lineage, training

Kimi K3 is a sparse mixture-of-experts model. Moonshot's own model list on the Kimi API Platform describes it as having "2.8 trillion parameters" with visual understanding and a 1M-token context window. That parameter count is the only architectural number this page could confirm from a Moonshot-controlled page.

The figures circulating elsewhere — that it routes the top 16 of 896 experts per token, that it is built on Kimi Delta Attention and Attention Residuals — appear in press coverage and third-party write-ups, not on any Moonshot documentation page read for this article. They are plausible and consistent, but they are not confirmed here, so they are not stated here as fact.

The open-weight question matters more than the architecture, and it has a clear answer today. As of 2026-07-25 there is no Kimi-K3 repository on the moonshotai Hugging Face organisation. The newest repositories there are Kimi-K2.7-Code, Kimi-K2.6 and Kimi-K2.5. So while K3 is widely described as an open-weight release, the weights were not published at the time of writing and no licence has been announced on a Moonshot page. Treat "open-weight" as a stated intention, not a present capability.

What is unambiguous is the serving behaviour, because Moonshot documents it precisely. K3 always runs with thinking enabled — the docs call it "Preserved Thinking" — and the only control over how hard it thinks is a top-level reasoning_effort field taking low, high or max, defaulting to max. The generation parameters that other models expose are fixed on the server side rather than exposed to you.

At a glanceSee it

Kimi K3 diagram

How a Kimi K3 request is priced and shaped, from cached prefix through always-on thinking to billed output.

CapacityContext, output and what fits

FactValueSource
Model idkimi-k3Kimi K3 quickstart
Context window1M tokens (1,048,576)Kimi K3 quickstart — "a 1M-token context window". Note: the K3 pricing page states no context figure
Max output tokensDefaults to 131,072; configurable up to 1,048,576 via max_completion_tokensKimi K3 quickstart
Input modalitiesText, images, videoConfigure Kimi Vision Models — "Kimi K3, Kimi K2.7 Code, and Kimi K2.6 all support text, image, and video input"
Output modalitiesText, plus a separate reasoning_content streamKimi K3 quickstart
Image and video input methodInline base64 data URI (data:image/png;base64,…). The docs recommend file upload instead "for larger videos, or images and videos that need to be referenced multiple times". Stated limits: images no larger than 4K resolution, videos no larger than 1080p. An ms://<file-id> URI scheme is not documented and could not be confirmed; neither could an explicit statement that public URLs are rejectedConfigure Kimi Vision Models
Knowledge cutoffNot disclosedNot stated on any Moonshot page read
Weights availableNot yet published, but announced. No Kimi-K3 repository exists on the moonshotai Hugging Face organisation as of 2026-07-26 — the newest published repo is Kimi-K2.7-Code. However the quickstart states: "The full model weights will be released by July 27, 2026"huggingface.co/moonshotai; Kimi K3 quickstart
LicenceNot disclosed. No licence is published on any Moonshot page read; none can be established until the weights are posted—
Release dateCould not be confirmed from a Moonshot-controlled page as of 2026-07-26—

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
messagesConversation so farRequiredTool definitions and prior turns all count against the 1M window and against input billing.
max_completion_tokensCaps generated tokensDefault 131,072; up to 1,048,576Raising it lets long agent turns finish; it does not make the model verbose. Prefer this name — max_tokens is still documented but is marked deprecated.
reasoning_effortHow much the model thinks before answeringlow, high, max. Default maxThe single biggest lever on both latency and bill. low cuts thinking tokens sharply; max is the default and the expensive one.
toolsFunction tools the model may callArray, default noneThe schema text is counted in input tokens, so a large tool set is a standing cost on every turn.
tool_choiceForces or forbids tool useauto (default), none, required, or a named functionrequired guarantees a call but can produce a call with invented arguments if no tool fits.
response_formatOutput shape{"type":"text"} default; json_schema (with strict) documented. Support for json_object on kimi-k3 could not be confirmed from the chat-completions referencejson_schema is the documented and reliable option when a downstream parser is involved.
stopHalts on a full matchString or array of strings, maximum 5 sequences, each ≤ 32 bytes; default nullUseful for delimiting agent scratchpads, but budget them — the cap is 5, not DeepSeek's 16. It does not stop the thinking stream.
streamServer-sent eventsDefault falseWith thinking always on, streaming is close to mandatory or the first token is far away.
stream_optionsAdds usage to the final chunk{"include_usage": true}The only way to see cache-hit versus cache-miss token split per request.
logprobsReturn log probabilitiesDefault falseUseful for confidence gating, since you cannot lower temperature to sharpen the distribution.
top_logprobsAlternatives per positionInteger, 0–20Costs response size, not tokens.
prompt_cache_keyGroups similar requests for cachingString, default nullHelps route repeated prefixes to the same cache; caching itself is automatic.
predictionPredicted output for regenerationObject (type, content), default nullSpeeds up edit-in-place patterns where most of the output is already known.
safety_identifierStable end-user identifierString, default nullUsed for abuse detection; has no effect on generation.
temperature, top_pSampling controlsNot supported on kimi-k3These belong to the Moonshot V1 request shape, not the K3 one. The quickstart states the server fixes temperature=1.0 and top_p=0.95 and advises omitting them. Steer with the prompt and response_format instead.
n, presence_penalty, frequency_penaltySample count and repetition penaltiesNot supported on kimi-k3Also Moonshot V1-only fields. Fixed at n=1 and both penalties at 0. For multiple candidates, issue multiple requests.

Two knobs actually matter. reasoning_effort is the whole steering wheel — it is the only thing that changes how the model deliberates, and since thinking is billed inside the $15 output rate, moving it from max to low is the most direct cost control you have. After that, response_format with a JSON schema is what buys back the determinism you lost when Moonshot took away temperature. Everything else is plumbing.

SamplingShaping the output distribution

There is nothing to tune. Moonshot's chat API reference lists temperature, top_p, n, presence_penalty and frequency_penalty as unsupported for kimi-k3, and the K3 quickstart states the server holds them at temperature=1.0, top_p=0.95, n=1 and both penalties at 0. That is a deliberate design: the model is tuned as a thinking model at one operating point, and Moonshot does not want callers moving off it. The practical consequence is that you cannot make K3 more deterministic by turning temperature down, and you cannot widen it for brainstorming either. Repeated identical requests will vary. If you need low-variance structured output, get it from response_format with a JSON schema rather than from sampling, and use logprobs if you need a confidence signal. If you genuinely need a temperature dial on a Moonshot model, the moonshot-v1-* generation still exposes one; K3 does not.

ReasoningThinking, effort and budgets

Thinking is always on and cannot be disabled. There is no thinking object on K3 the way there is on Kimi K2.6 and K2.5 — instead K3 takes a top-level reasoning_effort string accepting low, high or max, with max as the default. The docs describe the behaviour as "Preserved Thinking", and reasoning is visible: the streamed response carries separate reasoning_content and content deltas, and the non-streaming message exposes a reasoning_content attribute. There is no token-budget parameter — you get three coarse levels, not a numeric cap, which is a real limitation next to the numeric budgets Alibaba and Anthropic expose. Whether thinking tokens are billed at the output rate is not stated on the K3 pricing page; the honest assumption is that they are, since they are generated tokens, but that is inference and not a documented fact. Because max is the default, the expensive setting is what you get if you send nothing.

ToolsFunction calling and server tools

Function calling works and parallelism is explicit: the tool-calls guide states the model "can return multiple tool_calls at once" and notes it will tend to call independent tools in parallel, so your executor should fan out rather than loop. tool_choice accepts auto, none, required or a named function, and structured output is available through response_format including json_schema. There is no documented built-in server-side web search tool for K3 — Moonshot's own examples build search and crawl as ordinary user-defined tools with JSON Schema. Three failure modes are called out in the docs and worth pre-empting. First, every tool_call must be answered by a matching message with role=tool and the correct tool_call_id, or the next request errors. Second, the tools block itself is counted in your input tokens, so a fat tool registry is a per-turn tax. Third, when streaming, delta.tool_calls arrives only after delta.content finishes — parsers that assume interleaving will hang.

CostPrice, caching, batching, what drives the bill

List price from Moonshot's own Kimi K3 pricing page, in USD per million tokens:

ItemPrice per 1M tokens
Input, cache miss$3.00
Input, cache hit$0.30
Output$15.00

The page states "Prices exclude applicable taxes" and shows no context-length tiering and no volume discount. Three things drive the actual bill. Caching is automatic — "Context Caching is automatically enabled for all model requests", no parameter required — but a cache hit only happens if the previous request's prompt tokens exceeded 256, so short prompts never cache; the 10x discount is real and applies to the repeated prefix only. Second, there is no batch discount available here: Moonshot's Batch API prices inference at 60% of the standard rate but lists only kimi-k2.7-code, kimi-k2.6 and kimi-k2.5 as eligible, not kimi-k3. Third, output is 5x input and thinking is always on at max by default, so output dominates unless you lower reasoning_effort. On vision input, the K3 pricing page publishes no separate image or video token rate and no resolution-based billing formula — how multimodal input is metered could not be confirmed. The vision guide gives only operational guidance: images should be no larger than 4K resolution and videos no larger than 1080p.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Moonshot Kimi APIYesBase URL https://api.moonshot.ai/v1, model id kimi-k3, OpenAI-SDK compatible.
Amazon BedrockNoBedrock's Moonshot AI section lists Kimi K2.5 and Kimi K2 Thinking only.
Google Vertex AINoThe Vertex managed-API model list shows Kimi K2 Thinking only.
Microsoft FoundryNoFoundry's Moonshot AI models sold by Azure are Kimi-K2.7-Code, Kimi-K2.6 and Kimi-K2.5, all Preview.
Hugging Face InferenceUnconfirmedA moonshotai/Kimi-K3 repository now exists; whether any inference provider serves it was not checked.
Self-hostingYesWeights published on Hugging Face as moonshotai/Kimi-K3, read 2026-07-30.

StrengthsWhat it is good at

  • 1,048,576-token context with a documented path to 1M output tokens via max_completion_tokens — the input and output ceilings are the same number, which is unusual.
  • Cache-hit input at $0.30 per million against a $3.00 miss rate, applied automatically with no cache-control parameter to manage.
  • Parallel tool calls are documented behaviour, not an accident: the guide explicitly tells you to execute returned tool_calls concurrently.
  • Native video understanding alongside images — the vision guide lists mp4, mov and webm for K3, which most competitors at this tier do not accept.
  • Reasoning is visible: reasoning_content streams separately from content, so you can log or display the deliberation.

LimitsWhere it falls down

  • No sampling control at all. temperature, top_p, n and both penalties are unsupported, so you cannot tighten or widen the distribution for a given task.
  • Weights are not published. No Kimi-K3 repository exists on the moonshotai Hugging Face organisation as of 2026-07-25, and no licence has been announced, so "open-weight" is a promise rather than a fact you can act on.
  • No batch discount. The BatchJob pricing page lists K2.7-Code, K2.6 and K2.5 as eligible; K3 is not among them.
  • Thinking cannot be turned off, and the budget is three coarse levels rather than a token cap — awkward for latency-sensitive classification work.
  • Not on Bedrock, Vertex or Microsoft Foundry, so the vendor API is the only route, with the jurisdictional consequences that implies.

Against its neighboursHow it compares

Against DeepSeek V4 Pro, the comparison is stark on price: DeepSeek charges $1.32 in and $3.96 out per million at peak, and half that off-peak, against Kimi's $3.00 and $15.00 — roughly 3.8x cheaper on output at peak and 7.6x off-peak — for the same 1M context. DeepSeek also gives you temperature, top_p, a thinking toggle you can switch off, and MIT-licensed weights you can actually download. Kimi's claim is native vision and video, which DeepSeek V4 does not have at all, and a stronger reported coding profile. Against Claude Fable 5, Kimi is the budget option with a comparable context window, but Fable is available on Bedrock, Vertex and Foundry with the procurement and residency options that brings, while Kimi is vendor-API-only today.

Getting startedThe smallest call that works

code
POST https://api.moonshot.ai/v1/chat/completions
Authorization: Bearer $MOONSHOT_API_KEY
Content-Type: application/json

{
  "model": "kimi-k3",
  "messages": [
    {"role": "user", "content": "Introduce Kimi K3 in one sentence."}
  ]
}

Change reasoning_effort first — add "reasoning_effort": "low" and watch latency and output tokens fall, because the default is max. Then add "stream": true with "stream_options": {"include_usage": true} so you can see the cache-hit split. Do not bother adding temperature; it is not supported and will not do anything.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-12 against https://platform.kimi.ai/. re-read 2026-09-12 (research pass, primary source read today: https://platform.kimi.ai/ (raw HTML, server-rendered; prices present in visible text) — figures confirmed: input_per_m 3.0, output_per_m 15.0, cache_hit_input_per_m 0.3, context 1M ('a 1M-token context window'), max_output not stated on the page.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $3 in / $15 out per M, $0.3 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-26 (currency proposal src-2026-08-26-bb1c34; read of platform.kimi.ai — Kimi K3 $3/$15, the $0.30 cache hit and the 1M context window all unchanged. ⚑ THE PAGE NOW LISTS A MODEL THIS FILE DOES NOT CARRY: 'Kimi K2.7 Code' at $0.95/$4.00 with a $0.19 cache hit and a 256k window, beside the K2.6 row at the same $0.95/$4.00. A plausible cause of the hash move, recorded as an observation; whether it earns a row is an operator call.)

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page Kimi K3 is available on Amazon Bedrock with a 1M-token context window, making it callable through managed AWS endpoints.

    China frontier labs · 18 Sep 2026 · source

  • Updated this page Moonshot is reportedly negotiating with Microsoft, Amazon, and Google to host Kimi K3 on their clouds, which would let you use the model without going through Moonshot directly.

    Add a dated note that Moonshot is in talks with Microsoft, Amazon, and Google to offer Kimi K3 through their clouds, with a reported 30% revenue share, so the model's availability may expand beyond Moonshot's own API.

    China frontier labs · 26 Aug 2026 · source

RecentAuto-linked from the brief, not a rewrite of this page

Showing the 10 most recent references. 4 older were dropped — a reference ages, so this list does not grow forever.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning