Home › Frontier Models › Muse Spark 1.1
Model · Reference

Muse Spark 1.1

Meta-native apps; long agentic runs

In one line

No Llama 5 exists as of 2026-07-25 — Meta's newest open weights are still Llama 4, and its current frontier model, Muse Spark 1.1, ships closed.

Why this oneWhat it is actually for

You are here because a row somewhere said Meta shipped a 600B open-weight model with a five-million-token window, and you were about to plan a self-hosted deployment around it. Stop. As of 2026-07-25 there is no Llama 5 on Meta's blog, on Meta's developer site, or in the meta-llama organisation on Hugging Face, whose newest repository of any kind was created on 2025-04-28. The real decision underneath the question splits two ways. If you need weights you control, the last Llama you can actually download is Llama 4, and Mistral Large 3 is the stronger open-weight frontier option today. If you want Meta's current best model, it is Muse Spark 1.1, and it is closed.

What it isArchitecture, lineage, training

There is nothing to describe, because the model does not exist. Every check pointed the same way. llama.com now issues a 301 to developer.meta.com/ai/. Meta's AI blog carries no Llama 5 post — its 2026 model announcements are Muse Spark, Muse Spark 1.1, Muse Image and Muse Video. The Hugging Face models API for the meta-llama organisation returns thirteen Llama 4 repositories created between 2025-04-01 and 2025-04-04, then Llama-Guard-4-12B and the Prompt-Guard-2 pair in late April 2025, and nothing after that.

What Meta does have as open weights is the Llama 4 herd: Llama-4-Scout-17B-16E-Instruct and Llama-4-Maverick-17B-128E-Instruct. The organisation card describes them as natively multimodal mixture-of-experts models — Scout with 17B active parameters across 16 experts, Maverick with 17B active across 128 — under the Llama 4 Community License, a custom commercial licence rather than an OSI one. Both repositories are gated behind an access request, so their config.json would not load for me and I am not quoting hidden sizes, head counts or rope theta for them.

Meta's frontier line moved elsewhere. Muse Spark was announced on 2026-04-08 by Meta Superintelligence Labs and updated to 1.1 on 2026-07-09. Meta discloses no parameter count and no architecture for it. What it does publish: the model id is muse-spark-1.1, the context window is 1,048,576 tokens, inputs are text, image, video, audio and PDF, and output is text. Weights are not released; the April announcement says only that Meta hopes to open-source future versions. The lineage the site's row assumed — open weights carrying forward from Llama into a bigger Llama — is precisely the thing Meta stopped doing.

At a glanceSee it

Muse Spark 1.1 diagram

Meta's newest open weights are Llama 4, from April 2025; its newest frontier model is Muse Spark 1.1, announced July 2026 and served closed. There is no Llama 5.

CapacityContext, output and what fits

FactValueSource
A model named Llama 5Does not exist. Not on Meta's AI blog, not on developer.meta.com, not under meta-llama on Hugging FaceHugging Face models API for meta-llama and ai.meta.com/blog, both checked 2026-07-25
Meta's newest open-weight LLMsLlama 4 Scout 17B-16E and Llama 4 Maverick 17B-128E, repositories created 2025-04-01 to 2025-04-04Hugging Face models API, meta-llama org, sorted by creation date
Llama 4 licenceLlama 4 Community License Agreement, a custom commercial licence, not OSI-approvedhuggingface.co/meta-llama organisation card
Meta's current frontier modelMuse Spark, announced 2026-04-08; Muse Spark 1.1 announced 2026-07-09about.fb.com newsroom and ai.meta.com/blog
Context window, muse-spark-1.11,048,576 tokensMeta Model API models documentation
Max output tokensNot disclosed. max_completion_tokens is documented as model-dependent with no printed ceilingMeta Model API chat completion feature page
Modalities inText, image, video, audio, PDFMeta Model API models documentation
Modalities outTextMeta Model API models documentation
Knowledge cutoffNot disclosed on any Meta page read this sessionai.developer.meta.com docs and both Muse Spark announcements
Muse Spark weightsNot released. Meta states it hopes to open-source future versions of the modelabout.fb.com, Introducing Muse Spark
Muse Spark parameter count and architectureNot disclosedabout.fb.com and ai.meta.com, both Muse Spark posts

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
temperatureScales the logits before sampling0 to 2, default 1.0Toward 0 the model converges on its highest-probability continuation; above 1 it starts admitting low-probability tokens and long agentic runs drift
top_pNucleus cut-off on the sampling pool0 < top_p <= 1, default 1.0Below 1 the tail is truncated before temperature is applied; at the default no truncation happens at all. top_p: 0 returns HTTP 400
max_completion_tokensCeiling on generated tokensModel-dependent default; Meta prints no numberSet it low and long answers are cut mid-sentence rather than summarised; it is a hard stop, not a hint
frequency_penaltyPenalises tokens by how often they have already appeared-2 to 2, default 0Positive values suppress verbatim repetition; push past about 0.5 and the model starts avoiding necessary domain vocabulary
presence_penaltyPenalises tokens that have appeared at all-2 to 2, default 0Positive values widen topic coverage; on structured output it degrades key-name consistency
reasoning_effortSelects how much internal reasoning the model spendsUsable values on Muse Spark: minimal, low, medium, high, xhigh; default set by the model. none appears in the API-level enum but is not supported by Muse Spark and returns HTTP 400The single biggest lever on both latency and bill. Raising it lengthens the run. You cannot switch reasoning off — Meta states Muse Spark always reasons, so there is no zero-reasoning floor to fall back to
seedRequests reproducible samplingInteger, default nullMeta does not promise determinism, only that the seed is honoured as a best effort
nNumber of completions per requestDefault 1; only 1 is supportedAny value greater than 1 returns HTTP 400. To sample candidates you must issue separate requests
prompt_cache_retentionHow long a cached prefix is keptin_memory or 24h, default null24h pins a long system prompt across a working day so repeat calls bill cached-input rate instead of full rate
logprobsWould return per-token log probabilitiesNot supported on Muse SparkThe field is documented with a default of false, but Meta states that because Muse Spark is a reasoning model logprobs: true returns HTTP 400. There is no way to get token-level confidence out of this model
top_logprobsHow many alternatives to report per positionDocumented as 0 to 20, default null — but unreachable on this modelIt is only meaningful with logprobs true, and logprobs true is rejected on Muse Spark, so this field has no effect here
toolsFunction definitions the model may callNot stated on the pages publishedMeta's feature page names the field and defers to the API reference, which lists it in the request schema without printing its shape
tool_choiceForces, permits or forbids a tool callNot stated on the pages publishedSame situation as tools — the field is real, the allowed values are not documented on the reachable pages
response_formatRequests structured outputNot stated on the pages publishedNamed in the request schema; no JSON-schema example is printed, so validate the response yourself
stopStop sequencesNot supportedDocumented as unsupported on reasoning models; returns HTTP 400. Truncate on your own delimiter after the fact, or bound the run with max_completion_tokens
top_kFixed-size candidate poolNot documentedtop_k does not appear anywhere in Meta's request schema, and Meta does not list it among the fields that return HTTP 400 — its behaviour could not be confirmed either way. Use top_p, which is the only pool control this API documents

Three knobs decide almost everything here. reasoning_effort is the one that moves cost and latency by multiples, and it is where you should start tuning — but note the floor is minimal, not none, because Muse Spark always reasons. prompt_cache_retention set to 24h is the difference between paying $1.25 and $0.15 per million on a fixed system prompt. And max_completion_tokens is your only circuit breaker, because stop is rejected. Note also that logprobs, logit_bias, verbosity, prediction, web_search_options, modalities and audio all return 400 — a request copied verbatim from an OpenAI integration will often fail on one of them, and any evaluation harness that scores token-level confidence will not work against this model at all.

SamplingShaping the output distribution

Meta exposes a conventional two-knob sampler and says almost nothing about how to set it. temperature runs 0 to 2 with a default of 1.0; top_p runs above 0 to 1 with a default of 1.0. At those defaults there is no nucleus truncation at all, so the raw distribution is sampled at full temperature — a noticeably looser starting point than several competing APIs. The documentation does not state whether the two are applied jointly or whether one should be left alone, and it publishes no recommended per-task values, so treat any advice you have read on that as unverified for this model. A sensible starting point is temperature 1.0 with top_p lowered to about 0.9 for anything you intend to parse, and temperature nearer 0.2 for extraction and classification. There is no top_k and no logit_bias, so you cannot ban a token or hard-bound the candidate pool; the only shaping available is the nucleus. On long agentic runs, reasoning_effort changes output character far more than either sampling knob.

ReasoningThinking, effort and budgets

Yes, and it is a first-class request field rather than a prompt trick. reasoning_effort accepts six values — none, minimal, low, medium, high and xhigh — with the default chosen by the model rather than pinned in the docs. That is a finer ladder than most vendors publish, and xhigh in particular has no equivalent elsewhere. There is no numeric token budget: you pick a rung, not a count. On the consumer side the same capability surfaces as a Thinking mode in the Meta AI app and on meta.ai. What Meta does not document on any page I could reach is whether the reasoning trace is returned to the caller, whether it is redacted or summarised, and whether reasoning tokens are billed as output. Those three answers materially change your cost model, and none of them could be confirmed from a primary source as of 2026-07-25 — budget for reasoning tokens being billable until Meta says otherwise.

ToolsFunction calling and server tools

The Meta Model API is built to be swapped in under existing tooling. Meta's own overview states it is drop-in compatible with the OpenAI SDK, the Anthropic SDK, and agent CLIs including OpenCode and Claude Code, and that both a Responses-style API and a Chat Completions endpoint are offered; LangChain, LlamaIndex and the Vercel AI SDK are named as working clients. tools, tool_choice, parallel_tool_calls and response_format all appear in the documented request schema, but the reachable pages do not print their shapes or allowed values, so you should expect OpenAI-compatible semantics and verify against a live call rather than trust a spec. Server-side web search grounding exists and is billed separately per query. The failure modes are concrete and easy to hit: n above 1 is rejected, stop and top_k return HTTP 400, and web_search_options — the OpenAI-shaped way of configuring search — is also a 400, so grounding must be enabled Meta's way, not OpenAI's.

CostPrice, caching, batching, what drives the bill

Meta publishes one flat price for the API rather than per-model rates, which makes the arithmetic unusually simple.

ItemPrice
Input, per 1M tokens$1.25
Cached input, per 1M tokens$0.15
Output, per 1M tokens$4.25
Web search grounding$2.50 per 1,000 search queries

Figures are from the Meta Model API pricing and rate limits page, read 2026-07-25. The cache discount is the lever that matters: cached input is 12% of the uncached rate, and prompt_cache_retention set to 24h keeps a prefix warm long enough for that to bite on real traffic. No batch discount is documented. Whether reasoning tokens count as billable output is not stated anywhere I could read — assume they do. Rate limits are 60 requests and 2,000,000 tokens per minute on the free tier, 3,000 requests and 4,000,000 tokens per minute on paid. Llama 4 weights carry no per-token price at all; the cost there is GPUs, and Meta publishes no guidance on how many.

Where it runsSurfaces and availability

SurfaceAvailableNotes
Meta Model APIYesBase URL https://api.meta.ai/v1, bearer token auth, public preview. muse-spark-1.1 is the only model id in the catalogue
Hugging Face — Llama 4 weightsYesScout and Maverick under meta-llama, gated behind an access request
Hugging Face — Muse Spark weightsNoNot published. Meta says it hopes to open-source future versions
Self-hosting Llama 4YesUnder the Llama 4 Community License, once access is granted
Self-hosting Muse SparkNoNo weights exist to host
Amazon BedrockUnverifiedNo Meta-published statement about Muse Spark on Bedrock was found this session
Google Vertex AIUnverifiedNot named on any Meta page read this session
Microsoft Foundry or AzureUnverifiedNot named on any Meta page read this session

StrengthsWhat it is good at

  • Llama 4 Scout and Maverick weights remain downloadable and commercially licensable under the Llama 4 Community License, so the on-prem path Meta opened is still open — just frozen at April 2025.
  • The Meta Model API is documented as drop-in compatible with the OpenAI SDK, the Anthropic SDK, and agent CLIs including OpenCode and Claude Code, which makes evaluating it a base-URL change rather than a rewrite.
  • A 1,048,576-token window on muse-spark-1.1, with Meta describing the model as actively compacting that context during long agentic runs rather than merely accepting it.
  • Cached input at $0.15 against $1.25 uncached is a 88% discount, and prompt_cache_retention: "24h" is a longer published retention than most vendors offer.
  • reasoning_effort exposes six rungs including xhigh, a finer ladder than the three-or-four-level controls that are more common.

LimitsWhere it falls down

  • There is no Llama 5. Claims of 600B parameters and a 5M-token window trace to aggregator articles and press-release syndication, not to any Meta source; Meta's own April 2026 announcement was Muse Spark.
  • Muse Spark weights are closed and the API is in public preview, so the open-weight promise that made Llama interesting does not carry forward to Meta's current best model.
  • stop, top_k, logit_bias, verbosity, prediction, web_search_options, modalities and audio all return HTTP 400, so OpenAI-shaped requests frequently fail on a field you did not think about.
  • n is capped at 1 — no best-of-k sampling in a single call.
  • Meta discloses no parameter count, no architecture and no knowledge cutoff for Muse Spark, so you cannot reason about its behaviour from first principles at all.

Against its neighboursHow it compares

Against Mistral Large 3, this is not a close call if weights matter to you: Mistral Large 3 ships 675B parameters under Apache 2.0 with a published params.json, and costs $0.50 in and $1.50 out per million on the vendor API — cheaper on both sides than Meta's $1.25 and $4.25. Meta's advantages are the 1M window against Mistral's 256k, and a documented reasoning control that Mistral does not list for its Large model. Against GPT-5.6 Sol the comparison is closer in kind, since both are closed frontier models sold per token, and the decision comes down to price, window and whether your agent stack already speaks one dialect. What Meta no longer competes on is the thing this page was filed under. If your requirement was literally an open-weight frontier model you can run yourself, Meta has not had one since April 2025.

Getting startedThe smallest call that works

code
POST https://api.meta.ai/v1/chat/completions
Authorization: Bearer $MODEL_API_KEY
Content-Type: application/json

{
  "model": "muse-spark-1.1",
  "messages": [
    {"role": "user", "content": "What are three differences between TCP and UDP?"}
  ]
}

Change reasoning_effort first — it is the only field that moves latency and cost by multiples, and the model picks its own default if you leave it out. Then add prompt_cache_retention set to "24h" once your system prompt stabilises. Do not add stop or top_k; both return HTTP 400. Bound long runs with max_completion_tokens instead.

SourcesWhere every claim above came from

  • AI at Meta Blog — post index, checked 2026-07-25; no Llama 5 entry, newest model posts are Muse Spark 1.1 and Muse Image / Muse Video
  • Introducing Muse Spark: Meta's Most Powerful Model Yet — announcement of 2026-04-08, states weights are not released
  • Introducing Muse Spark 1.1 — 1 million token context, public preview on the Meta Model API
  • Meta Model API — Overview — base URL, bearer auth, OpenAI and Anthropic SDK compatibility, muse-spark-1.1 as the only model
  • Meta Model API — Models — 1,048,576 token window, input and output modalities
  • Meta Model API — Chat completion — parameter names, defaults, ranges, and the list of fields that return HTTP 400
  • Meta Model API — Create a chat completion — endpoint, headers and the verbatim curl example reproduced above
  • Meta Model API — Pricing and rate limits — $1.25 / $0.15 / $4.25 per 1M tokens, web search at $2.50 per 1,000 queries, tier rate limits. Checked 2026-07-25 at ai.developer.meta.com/docs/getting-started/pricing-rate-limits; that page stopped resolving by 2026-08-04, so it is named rather than linked. The figures stand as verified on the date shown — a citation records where a number came from, and the source moving does not unmake that.
  • meta-llama on Hugging Face — organisation card and models API, sorted by creation date; newest repository created 2025-04-28, Llama 4 is the latest family
  • Could not confirm: any model named Llama 5, in any form, from any Meta source — the 600B / 5M-context / April 2026 figures in the site's row are not supported and should be treated as wrong. Also unconfirmed: Muse Spark's parameter count, architecture, knowledge cutoff and max output tokens; whether its reasoning traces are returned or billed; the schemas for tools, tool_choice and response_format; and Muse Spark availability on Bedrock, Vertex AI or Azure. Llama 4's config.json could not be read because the repositories are access-gated.
Checked

Price and capacity verified 2026-09-18 against https://dev.meta.ai/docs/pricing-rate-limits.md. re-read 2026-09-18 (primary source read today: https://dev.meta.ai/docs/pricing-rate-limits.md — the page lists NO Llama model of any version; it carries the Muse family, and this row's displayed name has correctly read 'Muse Spark 1.1' since the 2026-07-30 correction. $1.25 per million input and $4.25 per million output is UNCHANGED. muse-spark-1.3 is now the flagship at the same headline price. The SLUG is stale and is deliberately not renamed: 206 kits join on it and an id is a stable slug. Closes the 2026-09-03 id-drift reading and the ageing ticket dated 2026-08-26.)

What changedWhat changed here

RecentAuto-linked from the brief, not a rewrite of this page
  • Muse will apparently let you download its entire filesystem 24 Sep · The Verge AI

    Two developers say Meta's Muse agent can be coaxed into zipping and sharing its entire root filesystem, including Ubuntu system files and internal documentation, with very little prompting. If you're building on agent harnesses, this is a concrete reminder that filesystem and tool access need hard sandboxing, not prompt-level instructions.

  • Everything new coming to Meta&#8217;s AI agent Muse 23 Sep · TechCrunch AI

    Meta went all-in on its Muse agent at Connect, putting it on smart glasses and announcing a standalone gadget. For a builder, the interesting part is the surface area: an agent with its own email address and device presence is a different integration target than a chat API.

  • Amazon blocks Meta AI agent from shopping on its platform 23 Sep · US frontier labs

    Amazon blocked Meta's AI agent Muse from shopping on its platform, an early concrete case of a retailer refusing agent traffic. If you're designing an agent that acts on third-party sites, assume platform-level blocking is a real failure mode and plan for it in your architecture.

  • Meta patches Muse exploit that let attackers control the AI agent 22 Sep · The Verge AI

    Meta patched a zero-day in its Muse macOS app that let local code redirect transcription processing away from Meta's servers and take control of the agent. If you're shipping or integrating a desktop agent, this is a concrete reminder that local-code-to-agent escalation paths are a real attack surface, not a theoretical one.

  • Meta&#8217;s AI agent has been blocked from using Amazon.com 21 Sep · TechCrunch AI

    Amazon blocked Meta's Muse agent from shopping on its site, citing an unauthorized AI agent violating its Conditions of Use. If you're building agents that act on third-party sites on a user's behalf, platform terms — not just technical capability — are the binding constraint on what your agent can do.

  • Amazon blocks Meta’s Muse AI agent 21 Sep · The Verge AI

    Amazon is blocking Meta's Muse agent from shopping on its behalf, showing users a message that unauthorized AI agent access violates its conditions of use. This is the clearest signal yet that agent-driven commerce will be gated by platform terms — design your agent's checkout and account flows assuming per-site permission, not open browsing.

  • MCP was always a bad idea? 20 Sep · Simon Willison

    Simon Willison's rebuttal argues MCP still matters precisely when you don't want a fully-permissioned terminal agent: it gives you service allow-lists, auth that keeps API keys away from the agent, a user-facing connect UI, and audit logging. If you're deciding whether to wire an agent directly to APIs or through MCP, this frames the tradeoff as control and auditability rather than capability.

  • Meta expands Muse agent connections, launches Muse for Mac 20 Sep · US frontier labs

    Meta expanded Muse's agent connections and shipped a Muse for Mac app. A desktop agent with access to local apps and files is a different integration surface than a browser agent — relevant if you're deciding where your own agent should live.

  • Meta&#8217;s Muse hits Mac, letting the AI take actions on your computer 18 Sep · TechCrunch AI

    Meta's Muse agent is now on Mac, where it can work with your files and apps and take actions on your behalf. A consumer agent with local file access is the pattern to study for permissioning and action-scoping, since that's exactly where your own agent designs will get scrutinized.

  • Meta AI launches Muse personal agent, including apps for iPhone and Mac 17 Sep · US frontier labs

    Meta launched Muse, a personal AI agent with iPhone and Mac apps, and it hit number two on the US Apple App Store. If you're building consumer agents, this is the distribution bar you're now competing against on iOS and macOS.

Showing the 10 most recent references. 3 older were dropped — a reference ages, so this list does not grow forever.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning