Moonshot's 1M-context flagship: always thinking, and the only model on this board that rejects temperature and top_p outright.
Why this oneWhat it is actually for
You reach for Kimi K3 when the job is long-horizon coding or agent work, the input is enormous, and you want frontier-class output at a fifth of what the US flagships charge for output tokens. The 1M window means a repository or a full document corpus goes in without chunking, and the cache-hit rate of $0.30 per million makes a fixed system prompt or codebase preamble nearly free to re-send. You reach for something else if you need deterministic sampling — K3 does not accept temperature or top_p at all — if you need a batch discount, or if your compliance posture cannot route data to a Chinese API endpoint. There is no self-host escape hatch yet either.
What it isArchitecture, lineage, training
Kimi K3 is a sparse mixture-of-experts model. Moonshot's own model list on the Kimi API Platform describes it as having "2.8 trillion parameters" with visual understanding and a 1M-token context window. That parameter count is the only architectural number this page could confirm from a Moonshot-controlled page.
The figures circulating elsewhere — that it routes the top 16 of 896 experts per token, that it is built on Kimi Delta Attention and Attention Residuals — appear in press coverage and third-party write-ups, not on any Moonshot documentation page read for this article. They are plausible and consistent, but they are not confirmed here, so they are not stated here as fact.
The open-weight question matters more than the architecture, and it has a clear answer today. As of 2026-07-25 there is no Kimi-K3 repository on the moonshotai Hugging Face organisation. The newest repositories there are Kimi-K2.7-Code, Kimi-K2.6 and Kimi-K2.5. So while K3 is widely described as an open-weight release, the weights were not published at the time of writing and no licence has been announced on a Moonshot page. Treat "open-weight" as a stated intention, not a present capability.
What is unambiguous is the serving behaviour, because Moonshot documents it precisely. K3 always runs with thinking enabled — the docs call it "Preserved Thinking" — and the only control over how hard it thinks is a top-level reasoning_effort field taking low, high or max, defaulting to max. The generation parameters that other models expose are fixed on the server side rather than exposed to you.
At a glanceSee it
How a Kimi K3 request is priced and shaped, from cached prefix through always-on thinking to billed output.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| Model id | kimi-k3 | Kimi K3 quickstart |
| Context window | 1M tokens (1,048,576) | Kimi K3 quickstart — "a 1M-token context window". Note: the K3 pricing page states no context figure |
| Max output tokens | Defaults to 131,072; configurable up to 1,048,576 via max_completion_tokens | Kimi K3 quickstart |
| Input modalities | Text, images, video | Configure Kimi Vision Models — "Kimi K3, Kimi K2.7 Code, and Kimi K2.6 all support text, image, and video input" |
| Output modalities | Text, plus a separate reasoning_content stream | Kimi K3 quickstart |
| Image and video input method | Inline base64 data URI (data:image/png;base64,…). The docs recommend file upload instead "for larger videos, or images and videos that need to be referenced multiple times". Stated limits: images no larger than 4K resolution, videos no larger than 1080p. An ms://<file-id> URI scheme is not documented and could not be confirmed; neither could an explicit statement that public URLs are rejected | Configure Kimi Vision Models |
| Knowledge cutoff | Not disclosed | Not stated on any Moonshot page read |
| Weights available | Not yet published, but announced. No Kimi-K3 repository exists on the moonshotai Hugging Face organisation as of 2026-07-26 — the newest published repo is Kimi-K2.7-Code. However the quickstart states: "The full model weights will be released by July 27, 2026" | huggingface.co/moonshotai; Kimi K3 quickstart |
| Licence | Not disclosed. No licence is published on any Moonshot page read; none can be established until the weights are posted | — |
| Release date | Could not be confirmed from a Moonshot-controlled page as of 2026-07-26 | — |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
messages | Conversation so far | Required | Tool definitions and prior turns all count against the 1M window and against input billing. |
max_completion_tokens | Caps generated tokens | Default 131,072; up to 1,048,576 | Raising it lets long agent turns finish; it does not make the model verbose. Prefer this name — max_tokens is still documented but is marked deprecated. |
reasoning_effort | How much the model thinks before answering | low, high, max. Default max | The single biggest lever on both latency and bill. low cuts thinking tokens sharply; max is the default and the expensive one. |
tools | Function tools the model may call | Array, default none | The schema text is counted in input tokens, so a large tool set is a standing cost on every turn. |
tool_choice | Forces or forbids tool use | auto (default), none, required, or a named function | required guarantees a call but can produce a call with invented arguments if no tool fits. |
response_format | Output shape | {"type":"text"} default; json_schema (with strict) documented. Support for json_object on kimi-k3 could not be confirmed from the chat-completions reference | json_schema is the documented and reliable option when a downstream parser is involved. |
stop | Halts on a full match | String or array of strings, maximum 5 sequences, each ≤ 32 bytes; default null | Useful for delimiting agent scratchpads, but budget them — the cap is 5, not DeepSeek's 16. It does not stop the thinking stream. |
stream | Server-sent events | Default false | With thinking always on, streaming is close to mandatory or the first token is far away. |
stream_options | Adds usage to the final chunk | {"include_usage": true} | The only way to see cache-hit versus cache-miss token split per request. |
logprobs | Return log probabilities | Default false | Useful for confidence gating, since you cannot lower temperature to sharpen the distribution. |
top_logprobs | Alternatives per position | Integer, 0–20 | Costs response size, not tokens. |
prompt_cache_key | Groups similar requests for caching | String, default null | Helps route repeated prefixes to the same cache; caching itself is automatic. |
prediction | Predicted output for regeneration | Object (type, content), default null | Speeds up edit-in-place patterns where most of the output is already known. |
safety_identifier | Stable end-user identifier | String, default null | Used for abuse detection; has no effect on generation. |
temperature, top_p | Sampling controls | Not supported on kimi-k3 | These belong to the Moonshot V1 request shape, not the K3 one. The quickstart states the server fixes temperature=1.0 and top_p=0.95 and advises omitting them. Steer with the prompt and response_format instead. |
n, presence_penalty, frequency_penalty | Sample count and repetition penalties | Not supported on kimi-k3 | Also Moonshot V1-only fields. Fixed at n=1 and both penalties at 0. For multiple candidates, issue multiple requests. |
Two knobs actually matter. reasoning_effort is the whole steering wheel — it is the only thing that changes how the model deliberates, and since thinking is billed inside the $15 output rate, moving it from max to low is the most direct cost control you have. After that, response_format with a JSON schema is what buys back the determinism you lost when Moonshot took away temperature. Everything else is plumbing.
SamplingShaping the output distribution
There is nothing to tune. Moonshot's chat API reference lists temperature, top_p, n, presence_penalty and frequency_penalty as unsupported for kimi-k3, and the K3 quickstart states the server holds them at temperature=1.0, top_p=0.95, n=1 and both penalties at 0. That is a deliberate design: the model is tuned as a thinking model at one operating point, and Moonshot does not want callers moving off it. The practical consequence is that you cannot make K3 more deterministic by turning temperature down, and you cannot widen it for brainstorming either. Repeated identical requests will vary. If you need low-variance structured output, get it from response_format with a JSON schema rather than from sampling, and use logprobs if you need a confidence signal. If you genuinely need a temperature dial on a Moonshot model, the moonshot-v1-* generation still exposes one; K3 does not.
ReasoningThinking, effort and budgets
Thinking is always on and cannot be disabled. There is no thinking object on K3 the way there is on Kimi K2.6 and K2.5 — instead K3 takes a top-level reasoning_effort string accepting low, high or max, with max as the default. The docs describe the behaviour as "Preserved Thinking", and reasoning is visible: the streamed response carries separate reasoning_content and content deltas, and the non-streaming message exposes a reasoning_content attribute. There is no token-budget parameter — you get three coarse levels, not a numeric cap, which is a real limitation next to the numeric budgets Alibaba and Anthropic expose. Whether thinking tokens are billed at the output rate is not stated on the K3 pricing page; the honest assumption is that they are, since they are generated tokens, but that is inference and not a documented fact. Because max is the default, the expensive setting is what you get if you send nothing.
ToolsFunction calling and server tools
Function calling works and parallelism is explicit: the tool-calls guide states the model "can return multiple tool_calls at once" and notes it will tend to call independent tools in parallel, so your executor should fan out rather than loop. tool_choice accepts auto, none, required or a named function, and structured output is available through response_format including json_schema. There is no documented built-in server-side web search tool for K3 — Moonshot's own examples build search and crawl as ordinary user-defined tools with JSON Schema. Three failure modes are called out in the docs and worth pre-empting. First, every tool_call must be answered by a matching message with role=tool and the correct tool_call_id, or the next request errors. Second, the tools block itself is counted in your input tokens, so a fat tool registry is a per-turn tax. Third, when streaming, delta.tool_calls arrives only after delta.content finishes — parsers that assume interleaving will hang.
CostPrice, caching, batching, what drives the bill
List price from Moonshot's own Kimi K3 pricing page, in USD per million tokens:
| Item | Price per 1M tokens |
|---|---|
| Input, cache miss | $3.00 |
| Input, cache hit | $0.30 |
| Output | $15.00 |
The page states "Prices exclude applicable taxes" and shows no context-length tiering and no volume discount. Three things drive the actual bill. Caching is automatic — "Context Caching is automatically enabled for all model requests", no parameter required — but a cache hit only happens if the previous request's prompt tokens exceeded 256, so short prompts never cache; the 10x discount is real and applies to the repeated prefix only. Second, there is no batch discount available here: Moonshot's Batch API prices inference at 60% of the standard rate but lists only kimi-k2.7-code, kimi-k2.6 and kimi-k2.5 as eligible, not kimi-k3. Third, output is 5x input and thinking is always on at max by default, so output dominates unless you lower reasoning_effort. On vision input, the K3 pricing page publishes no separate image or video token rate and no resolution-based billing formula — how multimodal input is metered could not be confirmed. The vision guide gives only operational guidance: images should be no larger than 4K resolution and videos no larger than 1080p.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Moonshot Kimi API | Yes | Base URL https://api.moonshot.ai/v1, model id kimi-k3, OpenAI-SDK compatible. |
| Amazon Bedrock | No | Bedrock's Moonshot AI section lists Kimi K2.5 and Kimi K2 Thinking only. |
| Google Vertex AI | No | The Vertex managed-API model list shows Kimi K2 Thinking only. |
| Microsoft Foundry | No | Foundry's Moonshot AI models sold by Azure are Kimi-K2.7-Code, Kimi-K2.6 and Kimi-K2.5, all Preview. |
| Hugging Face Inference | Unconfirmed | A moonshotai/Kimi-K3 repository now exists; whether any inference provider serves it was not checked. |
| Self-hosting | Yes | Weights published on Hugging Face as moonshotai/Kimi-K3, read 2026-07-30. |
StrengthsWhat it is good at
- 1,048,576-token context with a documented path to 1M output tokens via
max_completion_tokens— the input and output ceilings are the same number, which is unusual. - Cache-hit input at $0.30 per million against a $3.00 miss rate, applied automatically with no cache-control parameter to manage.
- Parallel tool calls are documented behaviour, not an accident: the guide explicitly tells you to execute returned
tool_callsconcurrently. - Native video understanding alongside images — the vision guide lists mp4, mov and webm for K3, which most competitors at this tier do not accept.
- Reasoning is visible:
reasoning_contentstreams separately fromcontent, so you can log or display the deliberation.
LimitsWhere it falls down
- No sampling control at all.
temperature,top_p,nand both penalties are unsupported, so you cannot tighten or widen the distribution for a given task. - Weights are not published. No
Kimi-K3repository exists on themoonshotaiHugging Face organisation as of 2026-07-25, and no licence has been announced, so "open-weight" is a promise rather than a fact you can act on. - No batch discount. The BatchJob pricing page lists K2.7-Code, K2.6 and K2.5 as eligible; K3 is not among them.
- Thinking cannot be turned off, and the budget is three coarse levels rather than a token cap — awkward for latency-sensitive classification work.
- Not on Bedrock, Vertex or Microsoft Foundry, so the vendor API is the only route, with the jurisdictional consequences that implies.
Against its neighboursHow it compares
Against DeepSeek V4 Pro, the comparison is stark on price: DeepSeek charges $1.32 in and $3.96 out per million at peak, and half that off-peak, against Kimi's $3.00 and $15.00 — roughly 3.8x cheaper on output at peak and 7.6x off-peak — for the same 1M context. DeepSeek also gives you temperature, top_p, a thinking toggle you can switch off, and MIT-licensed weights you can actually download. Kimi's claim is native vision and video, which DeepSeek V4 does not have at all, and a stronger reported coding profile. Against Claude Fable 5, Kimi is the budget option with a comparable context window, but Fable is available on Bedrock, Vertex and Foundry with the procurement and residency options that brings, while Kimi is vendor-API-only today.
Getting startedThe smallest call that works
POST https://api.moonshot.ai/v1/chat/completions
Authorization: Bearer $MOONSHOT_API_KEY
Content-Type: application/json
{
"model": "kimi-k3",
"messages": [
{"role": "user", "content": "Introduce Kimi K3 in one sentence."}
]
}Change reasoning_effort first — add "reasoning_effort": "low" and watch latency and output tokens fall, because the default is max. Then add "stream": true with "stream_options": {"include_usage": true} so you can see the cache-hit split. Do not bother adding temperature; it is not supported and will not do anything.
SourcesWhere every claim above came from
- Kimi K3 — Kimi API Platform
- Flagship Model Kimi K3 Pricing — Kimi API Platform
- Create Chat Completion — Kimi API Platform
- Model List — Kimi API Platform
- Reasoning Effort — Kimi API Platform
- Tool Calls — Kimi API Platform
- Configure Kimi Vision Models — Kimi API Platform
- Context Caching — Kimi API Platform
- BatchJob Pricing — Kimi API Platform
- moonshotai — Hugging Face organisation
- moonshotai/Kimi-K3 — Hugging Face model card
- Models at a glance — Amazon Bedrock
- Foundry Models sold by Azure — Microsoft Foundry
- Could not confirm from a primary source as of 2026-07-25: the expert count and per-token routing (widely reported as 16 of 896), the Kimi Delta Attention and Attention Residuals architecture claims, the release date, the knowledge cutoff, whether thinking tokens are billed separately from output tokens, and any open-weight licence. None of these appear on a Moonshot-controlled page read for this article.
Price and capacity verified 2026-09-12 against https://platform.kimi.ai/. re-read 2026-09-12 (research pass, primary source read today: https://platform.kimi.ai/ (raw HTML, server-rendered; prices present in visible text) — figures confirmed: input_per_m 3.0, output_per_m 15.0, cache_hit_input_per_m 0.3, context 1M ('a 1M-token context window'), max_output not stated on the page.) Prior: re-read 2026-09-05 (P2; build/currency number-anchored watch, run 2026-09-05 — every figure this row quotes was found on the cited page today: $3 in / $15 out per M, $0.3 cache-hit input. number_missing is empty and the page digest did not move, so no proposal was raised. DATE RE-STAMP ONLY: no price, context, access or expiry field changed.) Prior: manual 2026-08-26 (currency proposal src-2026-08-26-bb1c34; read of platform.kimi.ai — Kimi K3 $3/$15, the $0.30 cache hit and the 1M context window all unchanged. ⚑ THE PAGE NOW LISTS A MODEL THIS FILE DOES NOT CARRY: 'Kimi K2.7 Code' at $0.95/$4.00 with a $0.19 cache hit and a 256k window, beside the K2.6 row at the same $0.95/$4.00. A plausible cause of the hash move, recorded as an observation; whether it earns a row is an operator call.)
What changedWhat changed here
Updated this page Kimi K3 is available on Amazon Bedrock with a 1M-token context window, making it callable through managed AWS endpoints.
Updated this page Moonshot is reportedly negotiating with Microsoft, Amazon, and Google to host Kimi K3 on their clouds, which would let you use the model without going through Moonshot directly.
Add a dated note that Moonshot is in talks with Microsoft, Amazon, and Google to offer Kimi K3 through their clouds, with a reported 30% revenue share, so the model's availability may expand beyond Moonshot's own API.
- Announcing Kimi K3 on Snowflake Cortex AI
Kimi K3 is now available in private preview on Snowflake Cortex AI, putting Moonshot's frontier model inside a mainstream enterprise data platform. If your stack already runs on Snowflake, this is the shortest path to testing a Chinese frontier model without standing up your own inference.
- Moonshot’s Kimi K3 lands on Amazon in key test for Chinese open-source AI income
Moonshot's Kimi K3 is now available on Amazon, framed as a test of whether Chinese open-source models can earn revenue outside China. For a builder, it means a frontier-class Chinese open-weight model is reachable through mainstream cloud infrastructure rather than a self-hosted download.
- Moonshot AI’s Kimi K3 Arrives on Amazon Bedrock With 1M-Token Context
Kimi K3 is now available on Amazon Bedrock with a 1M-token context window, meaning you can call a Chinese frontier open-weight model through your existing AWS IAM, billing, and VPC setup instead of standing up your own inference. Worth testing as a drop-in for long-context workloads where you already run on Bedrock.
- Introducing Kimi K3 on Amazon Bedrock
AWS published the official Kimi K3 on Bedrock announcement, confirming the model is generally available through the platform rather than a rumor. For a builder this is the practical detail: model choice now spans US and Chinese frontier labs inside one cloud billing and IAM boundary.
- Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
Cognition released SWE-2, a Kimi K3 post-trained coding model that reportedly matches Fable 5.1 on FrontierCode at 64% lower cost. That's a concrete cost-per-quality datapoint for anyone choosing a coding agent backend — post-training a strong open base is now a competitive route, not just a research exercise.
- Moonshot AI eyes $2B annualized revenue as Kimi K3 lifts sales
Moonshot is reportedly eyeing $2B in annualized revenue on the back of Kimi K3 sales, and is weighing Hong Kong and mainland listings. It's a signal that the open-weight Chinese labs are now commercially durable enough to be a real second source for your model stack.
- Moonshot AI Seeks Up to 30% Revenue Share as Kimi K3 Eyes US Clouds
Moonshot AI is seeking up to a 30% revenue share as it negotiates to host Kimi K3 on US cloud infrastructure. If it lands, US builders could call a Chinese frontier model through domestic clouds, which changes the compliance and latency calculus for choosing Kimi K3.
- OpenAI-backed legal tech firm pivots to Chinese Kimi K3 open-weight model
An OpenAI-backed legal tech firm is pivoting to Moonshot's open-weight Kimi K3, a signal that even previously proprietary-model shops are switching to Chinese open-weight for cost or capability. For builders, it's another sign that open-weight is now a default path rather than a compromise.
- Reports indicate that Microsoft, Amazon, and Google are in discussions to launch services using the high-performance Chinese-made 'Kimi K3' robot, with developer Moonshot AI proposing a 30% revenue share.
Moonshot is in talks with Microsoft, Amazon, and Google to offer Kimi K3 through their clouds, asking for a 30% revenue share — if it closes, a Chinese frontier model becomes a one-click API on the US clouds you already build on.
Showing the 10 most recent references. 4 older were dropped — a reference ages, so this list does not grow forever.
Three kinds of claim, strongest first. Signal runs every morning.