xAI's coding and agentic frontier model — 500K context, February 2026 knowledge cutoff, first-party X search, and zero architecture disclosure.
Why this oneWhat it is actually for
Grok 4.5 is the model to reach for when you want frontier-tier coding and agentic work from a model whose pitch is finishing quickly rather than thinking exhaustively. xAI's announcement leans on output efficiency — it claims Grok 4.5 resolves tasks in 15,954 output tokens on average — and on serving speed of about 80 tokens per second. Two other things are genuinely differentiated. Its knowledge cutoff is 1 February 2026, months fresher than the Gemini 3 family's January 2025. And it is the one frontier model with first-party x_search as a server-side tool, which matters if your product has to reason over what is being posted right now rather than what was indexed.
What it isArchitecture, lineage, training
Grok 4.5 is a closed-weight reasoning model served only through xAI's own API. On architecture, xAI discloses nothing at all — not dense versus mixture-of-experts, not a parameter count, not layer or head counts, not a vocabulary size. The model page and the announcement are both silent. That is a material gap relative to Google, which at least names the family as sparse MoE, and it means every capacity question about this model has to be answered from behaviour rather than structure.
What xAI does describe is training. The announcement says Grok 4.5 was trained on datasets spanning coding, science, engineering and mathematics, across tens of thousands of NVIDIA GB300 GPUs, with training and stability techniques designed for large-scale runs. The reinforcement-learning stage is described as covering hundreds of thousands of tasks centred on multi-step software engineering and other technical work, graded automatically and by model-based graders. xAI also says the model was trained alongside Cursor — that is a vendor claim about a coding-agent partnership, not a published methodology.
The naming is worth understanding before you shop. xAI's model catalogue is not linearly versioned: grok-4.3, grok-4.20-0309-reasoning and grok-4.20-multi-agent-0309 all carry 1M-token windows, larger than Grok 4.5's 500K, at lower per-token prices. Grok 4.5 is positioned as the intelligence flagship, not the largest context. Its aliases are grok-4.5-latest and grok-build-latest. Reasoning is on by default at high effort, and the model card page lists function calling, structured outputs and reasoning as supported, with text and image inputs producing text output.
At a glanceSee it
How a Grok 4.5 request is priced and shaped, including the 200K billing threshold.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model string | grok-4.5, aliases grok-4.5-latest and grok-build-latest | xAI docs, Grok 4.5 model page |
| Context window | 500,000 tokens | xAI docs, Grok 4.5 model page and models table |
| Max output tokens | Not disclosed as a model limit; max_completion_tokens defaults to 128,000 | xAI REST API reference, chat completions |
| Input modalities | Text and images. Supported image types are jpg/jpeg and png | xAI docs, Grok 4.5 model page and image understanding guide |
| Output modalities | Text | xAI docs, Grok 4.5 model page |
| Knowledge cutoff | 1 February 2026 | xAI docs, Grok 4.5 model page |
| Released on the API | July 2026 | xAI docs release notes |
| Architecture | Not disclosed | No architecture statement on any xAI page read |
| Parameter count | Not disclosed | No parameter statement on any xAI page read |
| Training hardware | Tens of thousands of NVIDIA GB300 GPUs | xAI Grok 4.5 announcement |
| Weights available | No — closed API only | xAI docs list API access only |
| Rate limits | 150 requests per second; 50,000,000 tokens per minute | xAI docs, Grok 4.5 model page |
| Regions | us-east-1, us-west-2 | xAI docs, Grok 4.5 model page |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
temperature | Sampling randomness | 0 to 2; no default stated in the request schema | Accepted on the chat completions endpoint. xAI publishes no recommended value for Grok 4.5, so treat 1.0 as the working assumption and change it only with evals. |
top_p | Nucleus sampling threshold | No range or default stated in the reference | Documented as accepted, with no guidance. Use one of temperature or top_p, not both. |
max_completion_tokens | Upper bound on the number of tokens generated for a completion | Defaults to 128,000 when unset | The default is generous, so runaway agentic loops are expensive by default. Set it low for classification work. |
max_tokens | Older output cap | Deprecated in favour of max_completion_tokens | Use max_completion_tokens instead. |
stop | Sequences that halt generation | Up to 4 sequences | Documented as not supported by reasoning models. The reference does not say what happens if you send it anyway, so do not depend on either an error or a silent drop — terminate with a schema instead. |
seed | Requests deterministic sampling | Integer | Best-effort: repeated requests with the same seed and parameters should return the same result. |
frequency_penalty | Penalises tokens by existing frequency | -2.0 to 2.0 | Documented as not supported by reasoning models. |
presence_penalty | Penalises tokens already present | -2.0 to 2.0 | Documented as not supported by grok-3 and reasoning models. |
logit_bias | Per-token bias map | -100 to 100 | Marked (Unsupported) in the reference. There is no token-level steering on this model. |
logprobs / top_logprobs | Return token log-probabilities | Boolean; integer 0 to 8 | The reference marks both as not supported by grok-4.20 and newer models, which includes Grok 4.5. It does not say whether they error or are ignored, so do not build confidence gating on them. |
reasoning_effort | How much the model thinks before answering | low, medium, high; default high — per the Grok 4.5 model page and the July release note. Note the generic request reference disagrees, listing a none value, a low default, and support limited to grok-4.3; the model-specific pages are the better source here | The main quality-latency-cost dial. On the Responses API the same control is the nested reasoning object with an effort field. |
response_format | Structured output | text, json_object, json_schema | With json_schema the response is constrained to your schema, within documented limits. |
tools / tool_choice | Declares callable functions and controls selection | Max 128 functions | The chat completions reference states only functions are supported in this array. xAI's server-side tools (web search, X search, code execution and others) are documented separately and are surfaced through the Responses API and search_parameters — do not assume a bare {"type": "web_search"} entry works on chat completions without testing. |
parallel_tool_calls | Allows several tool calls in one turn | Boolean; set false to limit the model to one tool call | Turn off when your tool handlers are not idempotent or share state. |
prompt_cache_key | Stable key for sticky routing and cache hits | String | This is the Responses API control. On Chat Completions the documented equivalent is the x-grok-conv-id header. xAI says to set one of them: it routes a conversation's requests to the same server, and without it you often pay full input price on a cache-cold server. |
service_tier | Scheduling tier | default or priority | Priority changes scheduling and billing; check the pricing page before enabling it in production. |
Three knobs decide how this model behaves. reasoning_effort defaults to high, so an untuned call is doing maximum thinking and billing you for it — dropping to low is the first move on latency-sensitive paths. Cache stickiness is the one people skip and then wonder why caching never lands: set x-grok-conv-id on Chat Completions, or prompt_cache_key on the Responses API. And the negative space matters as much: stop, presence_penalty, frequency_penalty, logit_bias and logprobs are all documented as unsupported here, so any porting layer that passes them through from an OpenAI-shaped client is sending fields the model will not honour. Other accepted fields include n, deferred, search_parameters, stream and stream_options.include_usage.
SamplingShaping the output distribution
xAI is much quieter than Google about sampling. The chat completions schema lists temperature with a 0 to 2 range and top_p with no stated range, and neither carries a documented default or a recommended value for Grok 4.5. What the docs do tell you is what you cannot do: frequency_penalty, presence_penalty and stop all error on reasoning models, and logit_bias is unsupported outright. So the usual toolkit for tightening a distribution is largely absent, and the shaping has to happen elsewhere — through reasoning_effort, which changes how much deliberation precedes the first token, and through response_format with a JSON schema, which xAI states guarantees conformance within its documented limits. A sensible starting point is to leave temperature and top_p unset entirely, set reasoning_effort explicitly rather than inheriting high, and constrain shape with a schema. Then A/B temperature only if you have an eval that can detect the difference.
ReasoningThinking, effort and budgets
Grok 4.5 is a reasoning model and thinking is on by default at high effort. Effort takes low, medium or high. The naming differs by surface: on chat completions it is the flat reasoning_effort field, and on the Responses API it is nested as reasoning with an effort key. Note one inconsistency in xAI's own documentation — the chat completions REST reference annotates reasoning_effort as applying to grok-4.3 only and shows a default of low, while the Grok 4.5 model page and the launch announcement both state configurable low, medium or high with a default of high. Trust the model page for Grok 4.5. Reasoning tokens are billed as part of total consumption and surface in completion_tokens_details.reasoning_tokens. The raw chain is not returned; you get summarisations, streamed via response.reasoning_summary_text.delta, or encrypted content requested through include: ["reasoning.encrypted_content"].
ToolsFunction calling and server tools
Function calling is supported, with parallel_tool_calls for independent calls and a 128-tool ceiling on the tools array. The distinctive part is the server-side set: web_search, x_search, code_interpreter and a collections search over uploaded documents, all declared in the same tools array as your own functions using a type string such as {"type": "web_search"}. The web search tool takes allowed_domains and excluded_domains, each capped at five, plus enable_image_understanding and enable_image_search. Structured output is response_format with type: "json_schema", and strict is implicitly true for tool calling. The schema subset is where people get caught: maxLength is only guaranteed to 2,048 characters and maxItems to 256, beyond which conformance falls back to model behaviour rather than enforcement; keywords such as not, if/then/else and multi-subschema allOf are best-effort only; and some patterns are rejected with a 400.
CostPrice, caching, batching, what drives the bill
Prices are from the xAI docs models table and the Grok 4.5 model page, read on 2026-07-26. US dollars per million tokens. The billing rule that dominates everything else is stated plainly on the models page: models listed with two rows use long-context pricing, and requests whose prompt reaches the listed token threshold are billed at the higher rate for all tokens in the request. Cross 200K input tokens and the whole call, output included, doubles.
| Line | Prompt under 200K | Prompt at or over 200K |
|---|---|---|
| Input | 2.00 | 4.00 |
| Cached input | 0.30 | 0.60 |
| Output | 6.00 | 12.00 |
Cached input is 15 percent of the standard rate, so roughly an 85 percent saving on hit — that is the ratio of the published rates, not a discount xAI quotes; the docs say only that cached tokens are billed at a substantially lower rate. It is not automatic either. xAI tells you to set a stable conversation key so requests route to the same server and hit a warm cache: the x-grok-conv-id header on Chat Completions, or prompt_cache_key on the Responses API. Without one you often pay full input price on a cache-cold server. Reasoning tokens are billed as output, and with reasoning_effort defaulting to high that is the line item to watch. Server-side tools are billed on two components, token usage plus tool invocations, and only successful invocations are charged. Usage responses carry cost_in_usd_ticks, where 1 USD equals 10,000,000,000 ticks, plus prompt_tokens_details.cached_tokens on Chat Completions or input_tokens_details.cached_tokens on the Responses API for verifying cache hits. xAI's counter-argument to the high output price is token efficiency: the announcement claims about 15,954 output tokens to resolve a SWE Bench Pro task on average.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| xAI Responses API | Yes | POST https://api.x.ai/v1/responses with Authorization: Bearer. The surface xAI documents first. |
| xAI Chat Completions API | Yes | POST https://api.x.ai/v1/chat/completions, the surface whose full request schema is published. |
| OpenAI-compatible SDKs | Yes | Point an OpenAI client at base URL https://api.x.ai/v1. |
| Microsoft Foundry / Azure AI Foundry | Unverified | The catalogue lists Grok 4, Grok 4 Fast, Grok 3, Grok 3 Mini, Grok Code Fast 1 and Grok 4.3. Grok 4.5 could not be confirmed there as of 2026-07-25. |
| Amazon Bedrock | Unverified | Not checked against a primary Bedrock source in this session. |
| Google Vertex AI | Unverified | Not checked against a primary Vertex source in this session. |
| Hugging Face Inference | No | Closed weights; no Hugging Face model card exists. |
| Self-hosting | No | Weights are not released. |
| Fine-tuning | Unverified | No tuning offering for Grok 4.5 appeared on any xAI page read. |
StrengthsWhat it is good at
- Knowledge cutoff of 1 February 2026, per the Grok 4.5 model page — a year fresher than the Gemini 3 family's January 2025, which reduces how much you need to ground.
- First-party
x_searchas a server-side tool alongsideweb_searchandcode_interpreter, all declarable in the sametoolsarray as your own functions. - Output token efficiency is the explicit design goal — xAI's announcement claims tasks resolved in 15,954 output tokens on average, which matters when output costs 6.00 per million.
- Published rate limits are high for a frontier model: 150 requests per second and 50 million tokens per minute on the model page.
- Structured outputs are strict by construction, with an explicitly documented JSON Schema subset and enforcement thresholds rather than vague best-effort language.
LimitsWhere it falls down
- Zero architecture disclosure. No dense-versus-MoE statement, no parameter count, no expert count. You cannot reason about its capacity or scaling behaviour from anything published.
- The 200K threshold re-prices the entire request, so a 210K-token prompt costs double per token on input and output alike — and the window is 500K, so half of it is expensive territory.
stop,presence_penaltyandfrequency_penaltyall error on reasoning models, andlogit_biasis unsupported — an OpenAI-shaped client will need its request builder changed, not just its base URL.logprobsandtop_logprobsare documented as silently ignored on recent models. Silent, not erroring, which is the worst failure mode for anything that gates on confidence.- xAI's own docs disagree with themselves on
reasoning_effort— the REST reference scopes it to grok-4.3 with alowdefault while the model page says high on Grok 4.5. Set it explicitly rather than relying on any documented default.
Against its neighboursHow it compares
Against Gemini 3.1 Pro the trade is context and modality for recency and speed. Gemini takes 1M tokens plus audio and video natively; Grok takes 500K and text plus images. But Grok's cutoff is February 2026 against Gemini's January 2025, and both charge the same 2.00 input — Grok's output at 6.00 is half Gemini's 12.00. Against GPT-5.6 Sol and Claude Opus 5, Grok's specific claim is finishing in fewer tokens; xAI's announcement cites 15,954 output tokens per resolved task versus 67,020 for Claude Opus 4.8 at max effort, which is a vendor benchmark rather than an independent one. The real cost against all three is disclosure — no architecture at all, and a documentation set that contradicts itself on parameter defaults.
Getting startedThe smallest call that works
curl https://api.x.ai/v1/responses \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "grok-4.5",
"input": "Fix this function and explain the bug: function median(a){a.sort();return a[a.length/2]}"
}'Set reasoning effort explicitly on the next call rather than inheriting high — that is most of your latency and a large share of the bill. Add prompt_cache_key before you put anything with a stable system prompt into production, or the 0.30 cached rate will not land reliably. Do not carry over stop or the penalty parameters from an OpenAI client; they error here.
SourcesWhere every claim above came from
- Grok 4.5 — xAI Docs
- Models — xAI Docs
- Grok 4.5 guide — xAI Docs
- Chat — Inference API REST reference, xAI Docs
- Reasoning — xAI Docs
- Structured Outputs — xAI Docs
- Tools Overview — xAI Docs
- Web Search tool — xAI Docs
- Image Understanding — xAI Docs
- Quickstart — xAI Docs
- Release Notes — xAI Docs
- Introducing Grok 4.5 — xAI announcement
- Could not confirm from a primary source as of 2026-07-25: any architecture or parameter detail whatsoever; a model-level maximum output token limit distinct from the 128,000
max_completion_tokensdefault; the per-invocation prices forweb_search,x_searchandcode_interpreter; theservice_tier: prioritymultiplier; and availability on Microsoft Foundry, Amazon Bedrock or Vertex AI. The announcement page at x.ai/news/grok-4-5 returned HTTP 403 to direct fetch, so its figures — GB300 training hardware, 80 tokens per second, 15,954 average output tokens, and the Opus 4.8 comparison — were read through a search index summary of that page rather than the page itself, and should be treated as vendor claims at one remove. The roster row this page was built from listed cached input at 0.50; the xAI models table says 0.30 under 200K and 0.60 at or above it.
Price and capacity verified 2026-09-18 against https://docs.x.ai/developers/models/grok-4.5.md. re-read 2026-09-18 (primary source read today: https://docs.x.ai/docs/models — grok-4.5 still listed at $2.00 per million input and $6.00 per million output below 200k prompt tokens, UNCHANGED. It is superseded by grok-4.6, which is now its own row; this row is kept because 202 kits cite it and the comparison must stay complete. Closes the currency watch's 2026-09-03 id-drift reading and the ageing ticket dated 2026-08-18.)