Home › Frontier Models › Grok 4.5
Model · Reference

Grok 4.5

Real-time / social / news

In one line

xAI's coding and agentic frontier model — 500K context, February 2026 knowledge cutoff, first-party X search, and zero architecture disclosure.

Why this oneWhat it is actually for

Grok 4.5 is the model to reach for when you want frontier-tier coding and agentic work from a model whose pitch is finishing quickly rather than thinking exhaustively. xAI's announcement leans on output efficiency — it claims Grok 4.5 resolves tasks in 15,954 output tokens on average — and on serving speed of about 80 tokens per second. Two other things are genuinely differentiated. Its knowledge cutoff is 1 February 2026, months fresher than the Gemini 3 family's January 2025. And it is the one frontier model with first-party x_search as a server-side tool, which matters if your product has to reason over what is being posted right now rather than what was indexed.

What it isArchitecture, lineage, training

Grok 4.5 is a closed-weight reasoning model served only through xAI's own API. On architecture, xAI discloses nothing at all — not dense versus mixture-of-experts, not a parameter count, not layer or head counts, not a vocabulary size. The model page and the announcement are both silent. That is a material gap relative to Google, which at least names the family as sparse MoE, and it means every capacity question about this model has to be answered from behaviour rather than structure.

What xAI does describe is training. The announcement says Grok 4.5 was trained on datasets spanning coding, science, engineering and mathematics, across tens of thousands of NVIDIA GB300 GPUs, with training and stability techniques designed for large-scale runs. The reinforcement-learning stage is described as covering hundreds of thousands of tasks centred on multi-step software engineering and other technical work, graded automatically and by model-based graders. xAI also says the model was trained alongside Cursor — that is a vendor claim about a coding-agent partnership, not a published methodology.

The naming is worth understanding before you shop. xAI's model catalogue is not linearly versioned: grok-4.3, grok-4.20-0309-reasoning and grok-4.20-multi-agent-0309 all carry 1M-token windows, larger than Grok 4.5's 500K, at lower per-token prices. Grok 4.5 is positioned as the intelligence flagship, not the largest context. Its aliases are grok-4.5-latest and grok-build-latest. Reasoning is on by default at high effort, and the model card page lists function calling, structured outputs and reasoning as supported, with text and image inputs producing text output.

At a glanceSee it

Grok 4.5 diagram

How a Grok 4.5 request is priced and shaped, including the 200K billing threshold.

CapacityContext, output and what fits

FactValueSource
API model stringgrok-4.5, aliases grok-4.5-latest and grok-build-latestxAI docs, Grok 4.5 model page
Context window500,000 tokensxAI docs, Grok 4.5 model page and models table
Max output tokensNot disclosed as a model limit; max_completion_tokens defaults to 128,000xAI REST API reference, chat completions
Input modalitiesText and images. Supported image types are jpg/jpeg and pngxAI docs, Grok 4.5 model page and image understanding guide
Output modalitiesTextxAI docs, Grok 4.5 model page
Knowledge cutoff1 February 2026xAI docs, Grok 4.5 model page
Released on the APIJuly 2026xAI docs release notes
ArchitectureNot disclosedNo architecture statement on any xAI page read
Parameter countNot disclosedNo parameter statement on any xAI page read
Training hardwareTens of thousands of NVIDIA GB300 GPUsxAI Grok 4.5 announcement
Weights availableNo — closed API onlyxAI docs list API access only
Rate limits150 requests per second; 50,000,000 tokens per minutexAI docs, Grok 4.5 model page
Regionsus-east-1, us-west-2xAI docs, Grok 4.5 model page

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
temperatureSampling randomness0 to 2; no default stated in the request schemaAccepted on the chat completions endpoint. xAI publishes no recommended value for Grok 4.5, so treat 1.0 as the working assumption and change it only with evals.
top_pNucleus sampling thresholdNo range or default stated in the referenceDocumented as accepted, with no guidance. Use one of temperature or top_p, not both.
max_completion_tokensUpper bound on the number of tokens generated for a completionDefaults to 128,000 when unsetThe default is generous, so runaway agentic loops are expensive by default. Set it low for classification work.
max_tokensOlder output capDeprecated in favour of max_completion_tokensUse max_completion_tokens instead.
stopSequences that halt generationUp to 4 sequencesDocumented as not supported by reasoning models. The reference does not say what happens if you send it anyway, so do not depend on either an error or a silent drop — terminate with a schema instead.
seedRequests deterministic samplingIntegerBest-effort: repeated requests with the same seed and parameters should return the same result.
frequency_penaltyPenalises tokens by existing frequency-2.0 to 2.0Documented as not supported by reasoning models.
presence_penaltyPenalises tokens already present-2.0 to 2.0Documented as not supported by grok-3 and reasoning models.
logit_biasPer-token bias map-100 to 100Marked (Unsupported) in the reference. There is no token-level steering on this model.
logprobs / top_logprobsReturn token log-probabilitiesBoolean; integer 0 to 8The reference marks both as not supported by grok-4.20 and newer models, which includes Grok 4.5. It does not say whether they error or are ignored, so do not build confidence gating on them.
reasoning_effortHow much the model thinks before answeringlow, medium, high; default high — per the Grok 4.5 model page and the July release note. Note the generic request reference disagrees, listing a none value, a low default, and support limited to grok-4.3; the model-specific pages are the better source hereThe main quality-latency-cost dial. On the Responses API the same control is the nested reasoning object with an effort field.
response_formatStructured outputtext, json_object, json_schemaWith json_schema the response is constrained to your schema, within documented limits.
tools / tool_choiceDeclares callable functions and controls selectionMax 128 functionsThe chat completions reference states only functions are supported in this array. xAI's server-side tools (web search, X search, code execution and others) are documented separately and are surfaced through the Responses API and search_parameters — do not assume a bare {"type": "web_search"} entry works on chat completions without testing.
parallel_tool_callsAllows several tool calls in one turnBoolean; set false to limit the model to one tool callTurn off when your tool handlers are not idempotent or share state.
prompt_cache_keyStable key for sticky routing and cache hitsStringThis is the Responses API control. On Chat Completions the documented equivalent is the x-grok-conv-id header. xAI says to set one of them: it routes a conversation's requests to the same server, and without it you often pay full input price on a cache-cold server.
service_tierScheduling tierdefault or priorityPriority changes scheduling and billing; check the pricing page before enabling it in production.

Three knobs decide how this model behaves. reasoning_effort defaults to high, so an untuned call is doing maximum thinking and billing you for it — dropping to low is the first move on latency-sensitive paths. Cache stickiness is the one people skip and then wonder why caching never lands: set x-grok-conv-id on Chat Completions, or prompt_cache_key on the Responses API. And the negative space matters as much: stop, presence_penalty, frequency_penalty, logit_bias and logprobs are all documented as unsupported here, so any porting layer that passes them through from an OpenAI-shaped client is sending fields the model will not honour. Other accepted fields include n, deferred, search_parameters, stream and stream_options.include_usage.

SamplingShaping the output distribution

xAI is much quieter than Google about sampling. The chat completions schema lists temperature with a 0 to 2 range and top_p with no stated range, and neither carries a documented default or a recommended value for Grok 4.5. What the docs do tell you is what you cannot do: frequency_penalty, presence_penalty and stop all error on reasoning models, and logit_bias is unsupported outright. So the usual toolkit for tightening a distribution is largely absent, and the shaping has to happen elsewhere — through reasoning_effort, which changes how much deliberation precedes the first token, and through response_format with a JSON schema, which xAI states guarantees conformance within its documented limits. A sensible starting point is to leave temperature and top_p unset entirely, set reasoning_effort explicitly rather than inheriting high, and constrain shape with a schema. Then A/B temperature only if you have an eval that can detect the difference.

ReasoningThinking, effort and budgets

Grok 4.5 is a reasoning model and thinking is on by default at high effort. Effort takes low, medium or high. The naming differs by surface: on chat completions it is the flat reasoning_effort field, and on the Responses API it is nested as reasoning with an effort key. Note one inconsistency in xAI's own documentation — the chat completions REST reference annotates reasoning_effort as applying to grok-4.3 only and shows a default of low, while the Grok 4.5 model page and the launch announcement both state configurable low, medium or high with a default of high. Trust the model page for Grok 4.5. Reasoning tokens are billed as part of total consumption and surface in completion_tokens_details.reasoning_tokens. The raw chain is not returned; you get summarisations, streamed via response.reasoning_summary_text.delta, or encrypted content requested through include: ["reasoning.encrypted_content"].

ToolsFunction calling and server tools

Function calling is supported, with parallel_tool_calls for independent calls and a 128-tool ceiling on the tools array. The distinctive part is the server-side set: web_search, x_search, code_interpreter and a collections search over uploaded documents, all declared in the same tools array as your own functions using a type string such as {"type": "web_search"}. The web search tool takes allowed_domains and excluded_domains, each capped at five, plus enable_image_understanding and enable_image_search. Structured output is response_format with type: "json_schema", and strict is implicitly true for tool calling. The schema subset is where people get caught: maxLength is only guaranteed to 2,048 characters and maxItems to 256, beyond which conformance falls back to model behaviour rather than enforcement; keywords such as not, if/then/else and multi-subschema allOf are best-effort only; and some patterns are rejected with a 400.

CostPrice, caching, batching, what drives the bill

Prices are from the xAI docs models table and the Grok 4.5 model page, read on 2026-07-26. US dollars per million tokens. The billing rule that dominates everything else is stated plainly on the models page: models listed with two rows use long-context pricing, and requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request. Cross 200K input tokens and the whole call, output included, doubles.

LinePrompt under 200KPrompt at or over 200K
Input2.004.00
Cached input0.300.60
Output6.0012.00

Cached input is 15 percent of the standard rate, so roughly an 85 percent saving on hit — that is the ratio of the published rates, not a discount xAI quotes; the docs say only that cached tokens are billed at a substantially lower rate. It is not automatic either. xAI tells you to set a stable conversation key so requests route to the same server and hit a warm cache: the x-grok-conv-id header on Chat Completions, or prompt_cache_key on the Responses API. Without one you often pay full input price on a cache-cold server. Reasoning tokens are billed as output, and with reasoning_effort defaulting to high that is the line item to watch. Server-side tools are billed on two components, token usage plus tool invocations, and only successful invocations are charged. Usage responses carry cost_in_usd_ticks, where 1 USD equals 10,000,000,000 ticks, plus prompt_tokens_details.cached_tokens on Chat Completions or input_tokens_details.cached_tokens on the Responses API for verifying cache hits. xAI's counter-argument to the high output price is token efficiency: the announcement claims about 15,954 output tokens to resolve a SWE Bench Pro task on average.

Where it runsSurfaces and availability

SurfaceAvailableNotes
xAI Responses APIYesPOST https://api.x.ai/v1/responses with Authorization: Bearer. The surface xAI documents first.
xAI Chat Completions APIYesPOST https://api.x.ai/v1/chat/completions, the surface whose full request schema is published.
OpenAI-compatible SDKsYesPoint an OpenAI client at base URL https://api.x.ai/v1.
Microsoft Foundry / Azure AI FoundryUnverifiedThe catalogue lists Grok 4, Grok 4 Fast, Grok 3, Grok 3 Mini, Grok Code Fast 1 and Grok 4.3. Grok 4.5 could not be confirmed there as of 2026-07-25.
Amazon BedrockUnverifiedNot checked against a primary Bedrock source in this session.
Google Vertex AIUnverifiedNot checked against a primary Vertex source in this session.
Hugging Face InferenceNoClosed weights; no Hugging Face model card exists.
Self-hostingNoWeights are not released.
Fine-tuningUnverifiedNo tuning offering for Grok 4.5 appeared on any xAI page read.

StrengthsWhat it is good at

  • Knowledge cutoff of 1 February 2026, per the Grok 4.5 model page — a year fresher than the Gemini 3 family's January 2025, which reduces how much you need to ground.
  • First-party x_search as a server-side tool alongside web_search and code_interpreter, all declarable in the same tools array as your own functions.
  • Output token efficiency is the explicit design goal — xAI's announcement claims tasks resolved in 15,954 output tokens on average, which matters when output costs 6.00 per million.
  • Published rate limits are high for a frontier model: 150 requests per second and 50 million tokens per minute on the model page.
  • Structured outputs are strict by construction, with an explicitly documented JSON Schema subset and enforcement thresholds rather than vague best-effort language.

LimitsWhere it falls down

  • Zero architecture disclosure. No dense-versus-MoE statement, no parameter count, no expert count. You cannot reason about its capacity or scaling behaviour from anything published.
  • The 200K threshold re-prices the entire request, so a 210K-token prompt costs double per token on input and output alike — and the window is 500K, so half of it is expensive territory.
  • stop, presence_penalty and frequency_penalty all error on reasoning models, and logit_bias is unsupported — an OpenAI-shaped client will need its request builder changed, not just its base URL.
  • logprobs and top_logprobs are documented as silently ignored on recent models. Silent, not erroring, which is the worst failure mode for anything that gates on confidence.
  • xAI's own docs disagree with themselves on reasoning_effort — the REST reference scopes it to grok-4.3 with a low default while the model page says high on Grok 4.5. Set it explicitly rather than relying on any documented default.

Against its neighboursHow it compares

Against Gemini 3.1 Pro the trade is context and modality for recency and speed. Gemini takes 1M tokens plus audio and video natively; Grok takes 500K and text plus images. But Grok's cutoff is February 2026 against Gemini's January 2025, and both charge the same 2.00 input — Grok's output at 6.00 is half Gemini's 12.00. Against GPT-5.6 Sol and Claude Opus 5, Grok's specific claim is finishing in fewer tokens; xAI's announcement cites 15,954 output tokens per resolved task versus 67,020 for Claude Opus 4.8 at max effort, which is a vendor benchmark rather than an independent one. The real cost against all three is disclosure — no architecture at all, and a documentation set that contradicts itself on parameter defaults.

Getting startedThe smallest call that works

code
curl https://api.x.ai/v1/responses \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "grok-4.5",
    "input": "Fix this function and explain the bug: function median(a){a.sort();return a[a.length/2]}"
  }'

Set reasoning effort explicitly on the next call rather than inheriting high — that is most of your latency and a large share of the bill. Add prompt_cache_key before you put anything with a stable system prompt into production, or the 0.30 cached rate will not land reliably. Do not carry over stop or the penalty parameters from an OpenAI client; they error here.

SourcesWhere every claim above came from

Checked

Price and capacity verified 2026-09-18 against https://docs.x.ai/developers/models/grok-4.5.md. re-read 2026-09-18 (primary source read today: https://docs.x.ai/docs/models — grok-4.5 still listed at $2.00 per million input and $6.00 per million output below 200k prompt tokens, UNCHANGED. It is superseded by grok-4.6, which is now its own row; this row is kept because 202 kits cite it and the comparison must stay complete. Closes the currency watch's 2026-09-03 id-drift reading and the ageing ticket dated 2026-08-18.)

A living map of modern AI — kept current every morning