No Llama 5 exists as of 2026-07-25 — Meta's newest open weights are still Llama 4, and its current frontier model, Muse Spark 1.1, ships closed.
Why this oneWhat it is actually for
You are here because a row somewhere said Meta shipped a 600B open-weight model with a five-million-token window, and you were about to plan a self-hosted deployment around it. Stop. As of 2026-07-25 there is no Llama 5 on Meta's blog, on Meta's developer site, or in the meta-llama organisation on Hugging Face, whose newest repository of any kind was created on 2025-04-28. The real decision underneath the question splits two ways. If you need weights you control, the last Llama you can actually download is Llama 4, and Mistral Large 3 is the stronger open-weight frontier option today. If you want Meta's current best model, it is Muse Spark 1.1, and it is closed.
What it isArchitecture, lineage, training
There is nothing to describe, because the model does not exist. Every check pointed the same way. llama.com now issues a 301 to developer.meta.com/ai/. Meta's AI blog carries no Llama 5 post — its 2026 model announcements are Muse Spark, Muse Spark 1.1, Muse Image and Muse Video. The Hugging Face models API for the meta-llama organisation returns thirteen Llama 4 repositories created between 2025-04-01 and 2025-04-04, then Llama-Guard-4-12B and the Prompt-Guard-2 pair in late April 2025, and nothing after that.
What Meta does have as open weights is the Llama 4 herd: Llama-4-Scout-17B-16E-Instruct and Llama-4-Maverick-17B-128E-Instruct. The organisation card describes them as natively multimodal mixture-of-experts models — Scout with 17B active parameters across 16 experts, Maverick with 17B active across 128 — under the Llama 4 Community License, a custom commercial licence rather than an OSI one. Both repositories are gated behind an access request, so their config.json would not load for me and I am not quoting hidden sizes, head counts or rope theta for them.
Meta's frontier line moved elsewhere. Muse Spark was announced on 2026-04-08 by Meta Superintelligence Labs and updated to 1.1 on 2026-07-09. Meta discloses no parameter count and no architecture for it. What it does publish: the model id is muse-spark-1.1, the context window is 1,048,576 tokens, inputs are text, image, video, audio and PDF, and output is text. Weights are not released; the April announcement says only that Meta hopes to open-source future versions. The lineage the site's row assumed — open weights carrying forward from Llama into a bigger Llama — is precisely the thing Meta stopped doing.
At a glanceSee it
Meta's newest open weights are Llama 4, from April 2025; its newest frontier model is Muse Spark 1.1, announced July 2026 and served closed. There is no Llama 5.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| A model named Llama 5 | Does not exist. Not on Meta's AI blog, not on developer.meta.com, not under meta-llama on Hugging Face | Hugging Face models API for meta-llama and ai.meta.com/blog, both checked 2026-07-25 |
| Meta's newest open-weight LLMs | Llama 4 Scout 17B-16E and Llama 4 Maverick 17B-128E, repositories created 2025-04-01 to 2025-04-04 | Hugging Face models API, meta-llama org, sorted by creation date |
| Llama 4 licence | Llama 4 Community License Agreement, a custom commercial licence, not OSI-approved | huggingface.co/meta-llama organisation card |
| Meta's current frontier model | Muse Spark, announced 2026-04-08; Muse Spark 1.1 announced 2026-07-09 | about.fb.com newsroom and ai.meta.com/blog |
| Context window, muse-spark-1.1 | 1,048,576 tokens | Meta Model API models documentation |
| Max output tokens | Not disclosed. max_completion_tokens is documented as model-dependent with no printed ceiling | Meta Model API chat completion feature page |
| Modalities in | Text, image, video, audio, PDF | Meta Model API models documentation |
| Modalities out | Text | Meta Model API models documentation |
| Knowledge cutoff | Not disclosed on any Meta page read this session | ai.developer.meta.com docs and both Muse Spark announcements |
| Muse Spark weights | Not released. Meta states it hopes to open-source future versions of the model | about.fb.com, Introducing Muse Spark |
| Muse Spark parameter count and architecture | Not disclosed | about.fb.com and ai.meta.com, both Muse Spark posts |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
temperature | Scales the logits before sampling | 0 to 2, default 1.0 | Toward 0 the model converges on its highest-probability continuation; above 1 it starts admitting low-probability tokens and long agentic runs drift |
top_p | Nucleus cut-off on the sampling pool | 0 < top_p <= 1, default 1.0 | Below 1 the tail is truncated before temperature is applied; at the default no truncation happens at all. top_p: 0 returns HTTP 400 |
max_completion_tokens | Ceiling on generated tokens | Model-dependent default; Meta prints no number | Set it low and long answers are cut mid-sentence rather than summarised; it is a hard stop, not a hint |
frequency_penalty | Penalises tokens by how often they have already appeared | -2 to 2, default 0 | Positive values suppress verbatim repetition; push past about 0.5 and the model starts avoiding necessary domain vocabulary |
presence_penalty | Penalises tokens that have appeared at all | -2 to 2, default 0 | Positive values widen topic coverage; on structured output it degrades key-name consistency |
reasoning_effort | Selects how much internal reasoning the model spends | Usable values on Muse Spark: minimal, low, medium, high, xhigh; default set by the model. none appears in the API-level enum but is not supported by Muse Spark and returns HTTP 400 | The single biggest lever on both latency and bill. Raising it lengthens the run. You cannot switch reasoning off — Meta states Muse Spark always reasons, so there is no zero-reasoning floor to fall back to |
seed | Requests reproducible sampling | Integer, default null | Meta does not promise determinism, only that the seed is honoured as a best effort |
n | Number of completions per request | Default 1; only 1 is supported | Any value greater than 1 returns HTTP 400. To sample candidates you must issue separate requests |
prompt_cache_retention | How long a cached prefix is kept | in_memory or 24h, default null | 24h pins a long system prompt across a working day so repeat calls bill cached-input rate instead of full rate |
logprobs | Would return per-token log probabilities | Not supported on Muse Spark | The field is documented with a default of false, but Meta states that because Muse Spark is a reasoning model logprobs: true returns HTTP 400. There is no way to get token-level confidence out of this model |
top_logprobs | How many alternatives to report per position | Documented as 0 to 20, default null — but unreachable on this model | It is only meaningful with logprobs true, and logprobs true is rejected on Muse Spark, so this field has no effect here |
tools | Function definitions the model may call | Not stated on the pages published | Meta's feature page names the field and defers to the API reference, which lists it in the request schema without printing its shape |
tool_choice | Forces, permits or forbids a tool call | Not stated on the pages published | Same situation as tools — the field is real, the allowed values are not documented on the reachable pages |
response_format | Requests structured output | Not stated on the pages published | Named in the request schema; no JSON-schema example is printed, so validate the response yourself |
stop | Stop sequences | Not supported | Documented as unsupported on reasoning models; returns HTTP 400. Truncate on your own delimiter after the fact, or bound the run with max_completion_tokens |
top_k | Fixed-size candidate pool | Not documented | top_k does not appear anywhere in Meta's request schema, and Meta does not list it among the fields that return HTTP 400 — its behaviour could not be confirmed either way. Use top_p, which is the only pool control this API documents |
Three knobs decide almost everything here. reasoning_effort is the one that moves cost and latency by multiples, and it is where you should start tuning — but note the floor is minimal, not none, because Muse Spark always reasons. prompt_cache_retention set to 24h is the difference between paying $1.25 and $0.15 per million on a fixed system prompt. And max_completion_tokens is your only circuit breaker, because stop is rejected. Note also that logprobs, logit_bias, verbosity, prediction, web_search_options, modalities and audio all return 400 — a request copied verbatim from an OpenAI integration will often fail on one of them, and any evaluation harness that scores token-level confidence will not work against this model at all.
SamplingShaping the output distribution
Meta exposes a conventional two-knob sampler and says almost nothing about how to set it. temperature runs 0 to 2 with a default of 1.0; top_p runs above 0 to 1 with a default of 1.0. At those defaults there is no nucleus truncation at all, so the raw distribution is sampled at full temperature — a noticeably looser starting point than several competing APIs. The documentation does not state whether the two are applied jointly or whether one should be left alone, and it publishes no recommended per-task values, so treat any advice you have read on that as unverified for this model. A sensible starting point is temperature 1.0 with top_p lowered to about 0.9 for anything you intend to parse, and temperature nearer 0.2 for extraction and classification. There is no top_k and no logit_bias, so you cannot ban a token or hard-bound the candidate pool; the only shaping available is the nucleus. On long agentic runs, reasoning_effort changes output character far more than either sampling knob.
ReasoningThinking, effort and budgets
Yes, and it is a first-class request field rather than a prompt trick. reasoning_effort accepts six values — none, minimal, low, medium, high and xhigh — with the default chosen by the model rather than pinned in the docs. That is a finer ladder than most vendors publish, and xhigh in particular has no equivalent elsewhere. There is no numeric token budget: you pick a rung, not a count. On the consumer side the same capability surfaces as a Thinking mode in the Meta AI app and on meta.ai. What Meta does not document on any page I could reach is whether the reasoning trace is returned to the caller, whether it is redacted or summarised, and whether reasoning tokens are billed as output. Those three answers materially change your cost model, and none of them could be confirmed from a primary source as of 2026-07-25 — budget for reasoning tokens being billable until Meta says otherwise.
ToolsFunction calling and server tools
The Meta Model API is built to be swapped in under existing tooling. Meta's own overview states it is drop-in compatible with the OpenAI SDK, the Anthropic SDK, and agent CLIs including OpenCode and Claude Code, and that both a Responses-style API and a Chat Completions endpoint are offered; LangChain, LlamaIndex and the Vercel AI SDK are named as working clients. tools, tool_choice, parallel_tool_calls and response_format all appear in the documented request schema, but the reachable pages do not print their shapes or allowed values, so you should expect OpenAI-compatible semantics and verify against a live call rather than trust a spec. Server-side web search grounding exists and is billed separately per query. The failure modes are concrete and easy to hit: n above 1 is rejected, stop and top_k return HTTP 400, and web_search_options — the OpenAI-shaped way of configuring search — is also a 400, so grounding must be enabled Meta's way, not OpenAI's.
CostPrice, caching, batching, what drives the bill
Meta publishes one flat price for the API rather than per-model rates, which makes the arithmetic unusually simple.
| Item | Price |
|---|---|
| Input, per 1M tokens | $1.25 |
| Cached input, per 1M tokens | $0.15 |
| Output, per 1M tokens | $4.25 |
| Web search grounding | $2.50 per 1,000 search queries |
Figures are from the Meta Model API pricing and rate limits page, read 2026-07-25. The cache discount is the lever that matters: cached input is 12% of the uncached rate, and prompt_cache_retention set to 24h keeps a prefix warm long enough for that to bite on real traffic. No batch discount is documented. Whether reasoning tokens count as billable output is not stated anywhere I could read — assume they do. Rate limits are 60 requests and 2,000,000 tokens per minute on the free tier, 3,000 requests and 4,000,000 tokens per minute on paid. Llama 4 weights carry no per-token price at all; the cost there is GPUs, and Meta publishes no guidance on how many.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Meta Model API | Yes | Base URL https://api.meta.ai/v1, bearer token auth, public preview. muse-spark-1.1 is the only model id in the catalogue |
| Hugging Face — Llama 4 weights | Yes | Scout and Maverick under meta-llama, gated behind an access request |
| Hugging Face — Muse Spark weights | No | Not published. Meta says it hopes to open-source future versions |
| Self-hosting Llama 4 | Yes | Under the Llama 4 Community License, once access is granted |
| Self-hosting Muse Spark | No | No weights exist to host |
| Amazon Bedrock | Unverified | No Meta-published statement about Muse Spark on Bedrock was found this session |
| Google Vertex AI | Unverified | Not named on any Meta page read this session |
| Microsoft Foundry or Azure | Unverified | Not named on any Meta page read this session |
StrengthsWhat it is good at
- Llama 4 Scout and Maverick weights remain downloadable and commercially licensable under the Llama 4 Community License, so the on-prem path Meta opened is still open — just frozen at April 2025.
- The Meta Model API is documented as drop-in compatible with the OpenAI SDK, the Anthropic SDK, and agent CLIs including OpenCode and Claude Code, which makes evaluating it a base-URL change rather than a rewrite.
- A 1,048,576-token window on
muse-spark-1.1, with Meta describing the model as actively compacting that context during long agentic runs rather than merely accepting it. - Cached input at $0.15 against $1.25 uncached is a 88% discount, and
prompt_cache_retention: "24h"is a longer published retention than most vendors offer. reasoning_effortexposes six rungs includingxhigh, a finer ladder than the three-or-four-level controls that are more common.
LimitsWhere it falls down
- There is no Llama 5. Claims of 600B parameters and a 5M-token window trace to aggregator articles and press-release syndication, not to any Meta source; Meta's own April 2026 announcement was Muse Spark.
- Muse Spark weights are closed and the API is in public preview, so the open-weight promise that made Llama interesting does not carry forward to Meta's current best model.
stop,top_k,logit_bias,verbosity,prediction,web_search_options,modalitiesandaudioall return HTTP 400, so OpenAI-shaped requests frequently fail on a field you did not think about.nis capped at 1 — no best-of-k sampling in a single call.- Meta discloses no parameter count, no architecture and no knowledge cutoff for Muse Spark, so you cannot reason about its behaviour from first principles at all.
Against its neighboursHow it compares
Against Mistral Large 3, this is not a close call if weights matter to you: Mistral Large 3 ships 675B parameters under Apache 2.0 with a published params.json, and costs $0.50 in and $1.50 out per million on the vendor API — cheaper on both sides than Meta's $1.25 and $4.25. Meta's advantages are the 1M window against Mistral's 256k, and a documented reasoning control that Mistral does not list for its Large model. Against GPT-5.6 Sol the comparison is closer in kind, since both are closed frontier models sold per token, and the decision comes down to price, window and whether your agent stack already speaks one dialect. What Meta no longer competes on is the thing this page was filed under. If your requirement was literally an open-weight frontier model you can run yourself, Meta has not had one since April 2025.
Getting startedThe smallest call that works
POST https://api.meta.ai/v1/chat/completions
Authorization: Bearer $MODEL_API_KEY
Content-Type: application/json
{
"model": "muse-spark-1.1",
"messages": [
{"role": "user", "content": "What are three differences between TCP and UDP?"}
]
}Change reasoning_effort first — it is the only field that moves latency and cost by multiples, and the model picks its own default if you leave it out. Then add prompt_cache_retention set to "24h" once your system prompt stabilises. Do not add stop or top_k; both return HTTP 400. Bound long runs with max_completion_tokens instead.
SourcesWhere every claim above came from
- AI at Meta Blog — post index, checked 2026-07-25; no Llama 5 entry, newest model posts are Muse Spark 1.1 and Muse Image / Muse Video
- Introducing Muse Spark: Meta's Most Powerful Model Yet — announcement of 2026-04-08, states weights are not released
- Introducing Muse Spark 1.1 — 1 million token context, public preview on the Meta Model API
- Meta Model API — Overview — base URL, bearer auth, OpenAI and Anthropic SDK compatibility,
muse-spark-1.1as the only model - Meta Model API — Models — 1,048,576 token window, input and output modalities
- Meta Model API — Chat completion — parameter names, defaults, ranges, and the list of fields that return HTTP 400
- Meta Model API — Create a chat completion — endpoint, headers and the verbatim curl example reproduced above
- Meta Model API — Pricing and rate limits — $1.25 / $0.15 / $4.25 per 1M tokens, web search at $2.50 per 1,000 queries, tier rate limits. Checked 2026-07-25 at ai.developer.meta.com/docs/getting-started/pricing-rate-limits; that page stopped resolving by 2026-08-04, so it is named rather than linked. The figures stand as verified on the date shown — a citation records where a number came from, and the source moving does not unmake that.
- meta-llama on Hugging Face — organisation card and models API, sorted by creation date; newest repository created 2025-04-28, Llama 4 is the latest family
- Could not confirm: any model named Llama 5, in any form, from any Meta source — the 600B / 5M-context / April 2026 figures in the site's row are not supported and should be treated as wrong. Also unconfirmed: Muse Spark's parameter count, architecture, knowledge cutoff and max output tokens; whether its reasoning traces are returned or billed; the schemas for
tools,tool_choiceandresponse_format; and Muse Spark availability on Bedrock, Vertex AI or Azure. Llama 4'sconfig.jsoncould not be read because the repositories are access-gated.
Price and capacity verified 2026-09-18 against https://dev.meta.ai/docs/pricing-rate-limits.md. re-read 2026-09-18 (primary source read today: https://dev.meta.ai/docs/pricing-rate-limits.md — the page lists NO Llama model of any version; it carries the Muse family, and this row's displayed name has correctly read 'Muse Spark 1.1' since the 2026-07-30 correction. $1.25 per million input and $4.25 per million output is UNCHANGED. muse-spark-1.3 is now the flagship at the same headline price. The SLUG is stale and is deliberately not renamed: 206 kits join on it and an id is a stable slug. Closes the 2026-09-03 id-drift reading and the ageing ticket dated 2026-08-26.)
What changedWhat changed here
- Muse will apparently let you download its entire filesystem
Two developers say Meta's Muse agent can be coaxed into zipping and sharing its entire root filesystem, including Ubuntu system files and internal documentation, with very little prompting. If you're building on agent harnesses, this is a concrete reminder that filesystem and tool access need hard sandboxing, not prompt-level instructions.
- Everything new coming to Meta’s AI agent Muse
Meta went all-in on its Muse agent at Connect, putting it on smart glasses and announcing a standalone gadget. For a builder, the interesting part is the surface area: an agent with its own email address and device presence is a different integration target than a chat API.
- Amazon blocks Meta AI agent from shopping on its platform
Amazon blocked Meta's AI agent Muse from shopping on its platform, an early concrete case of a retailer refusing agent traffic. If you're designing an agent that acts on third-party sites, assume platform-level blocking is a real failure mode and plan for it in your architecture.
- Meta patches Muse exploit that let attackers control the AI agent
Meta patched a zero-day in its Muse macOS app that let local code redirect transcription processing away from Meta's servers and take control of the agent. If you're shipping or integrating a desktop agent, this is a concrete reminder that local-code-to-agent escalation paths are a real attack surface, not a theoretical one.
- Meta’s AI agent has been blocked from using Amazon.com
Amazon blocked Meta's Muse agent from shopping on its site, citing an unauthorized AI agent violating its Conditions of Use. If you're building agents that act on third-party sites on a user's behalf, platform terms — not just technical capability — are the binding constraint on what your agent can do.
- Amazon blocks Meta’s Muse AI agent
Amazon is blocking Meta's Muse agent from shopping on its behalf, showing users a message that unauthorized AI agent access violates its conditions of use. This is the clearest signal yet that agent-driven commerce will be gated by platform terms — design your agent's checkout and account flows assuming per-site permission, not open browsing.
- MCP was always a bad idea?
Simon Willison's rebuttal argues MCP still matters precisely when you don't want a fully-permissioned terminal agent: it gives you service allow-lists, auth that keeps API keys away from the agent, a user-facing connect UI, and audit logging. If you're deciding whether to wire an agent directly to APIs or through MCP, this frames the tradeoff as control and auditability rather than capability.
- Meta expands Muse agent connections, launches Muse for Mac
Meta expanded Muse's agent connections and shipped a Muse for Mac app. A desktop agent with access to local apps and files is a different integration surface than a browser agent — relevant if you're deciding where your own agent should live.
- Meta’s Muse hits Mac, letting the AI take actions on your computer
Meta's Muse agent is now on Mac, where it can work with your files and apps and take actions on your behalf. A consumer agent with local file access is the pattern to study for permissioning and action-scoping, since that's exactly where your own agent designs will get scrutinized.
- Meta AI launches Muse personal agent, including apps for iPhone and Mac
Meta launched Muse, a personal AI agent with iPhone and Mac apps, and it hit number two on the US Apple App Store. If you're building consumer agents, this is the distribution bar you're now competing against on iOS and macOS.
Showing the 10 most recent references. 3 older were dropped — a reference ages, so this list does not grow forever.
Three kinds of claim, strongest first. Signal runs every morning.