Home › Frontier Models › Jamba
Model · Reference

Jamba

Long-context at lower cost

In one line

AI21's hybrid Mamba-Transformer family: 256K context, mixture-of-experts, Apache 2.0 weights for Jamba2 Mini, and a hard 4096-token output cap on the API.

Why this oneWhat it is actually for

You reach for Jamba when the input is long, the budget is tight and the data cannot leave your network. The hybrid architecture puts a Mamba state-space block in most layers and full attention in only one layer of every eight, so the KV cache that normally makes 256K contexts expensive largely is not there - the state is fixed size. That is the whole pitch: long-context work at a cost curve that does not blow up with sequence length. AI21 sells this primarily as a private deployment - VPC, on-premises or hybrid - with the weights on Hugging Face, Jamba2 Mini under Apache 2.0. If you are doing document QA, classification or grounded retrieval over big corpora inside a regulated network, this is the shape that fits.

What it isArchitecture, lineage, training

Jamba is a joint attention and Mamba architecture - AI21's documentation gives the model type literally as "Joint Attention and Mamba (Jamba)". Most layers are Mamba state-space blocks; attention appears periodically. Mixture-of-experts is layered on top of that, with experts on alternating layers.

The real numbers are in config.json. For ai21labs/AI21-Jamba2-Mini: hidden_size 4096, num_hidden_layers 32, num_attention_heads 32, num_key_value_heads 8, vocab_size 65536, max_position_embeddings 262144. The interleaving is set by attn_layer_period 8 with attn_layer_offset 4 - one attention layer in every eight - and expert_layer_period 2 with expert_layer_offset 1, so every other layer is a mixture-of-experts layer. Those MoE layers carry num_experts 16 with num_experts_per_tok 2. The Mamba blocks are configured with mamba_d_state 16, mamba_d_conv 4, mamba_expand 2 and mamba_dt_rank 256. There is no rope_theta key at all, which is the tell that attention here does not use rotary position embeddings - position information comes from the recurrent blocks.

The Hugging Face safetensors index reports 51,570,323,328 parameters in BF16 for Jamba2 Mini, matching AI21's "52B parameters, 12B active" claim. Jamba Large is stated at 398B total and 94B active. ai21labs/AI21-Jamba2-3B, by contrast, is dense - its config.json sets num_experts to 1 - with hidden_size 2560, 28 layers, 20 attention heads, a single KV head and attn_layer_period 14.

AI21 describes Jamba2 training as starting from Jamba 1.5 pre-training, then mid-training on 500B curated tokens weighted toward maths, code, high-quality web text and long documents, a "state passing" phase that tunes the Mamba layers for context-length generalisation, cold-start supervised fine-tuning, DPO, then several rounds of on-policy reinforcement learning moving from short-context verifiable rewards to long-context mixed rewards.

At a glanceSee it

Jamba diagram

Mamba blocks carry the long context in fixed state while sparse experts and rare attention layers do the work.

CapacityContext, output and what fits

FactValueSource
API model stringsjamba-large currently points to jamba-large-1.7-2025-07; jamba-mini currently points to jamba-mini-2-2026-01AI21 Jamba foundation models page, API Versioning
Context window256K. max_position_embeddings is 262144 in config.json.AI21 docs; AI21-Jamba2-Mini config.json
Max output tokens4096. "For Jamba models, the maximum allowed value is 4096 tokens."AI21 Chat request reference, max_tokens
Model sizesJamba Large 398B total / 94B active; Jamba Mini 52B total / 12B active; Jamba 3B dense, no API endpointAI21 Jamba foundation models, Model Details table
Measured parameter count51,570,323,328 BF16 for AI21-Jamba2-MiniHugging Face safetensors index
Modalities in / outText in, text outAI21 Jamba foundation models, Model Details
Knowledge cutoffAugust 22nd, 2024 per the Model Details block. The Limitations block on the same page says the model "was trained on a dataset created in March 2024" - AI21's own page disagrees with itself.AI21 Jamba foundation models
Languages9 officially: English, Spanish, French, Portuguese, Italian, Dutch, German, Arabic, HebrewAI21 Jamba foundation models
Weights availableYes. ai21labs/AI21-Jamba2-Mini under Apache 2.0. ai21labs/AI21-Jamba-Large-1.7 under the Jamba Open Model License and gated on Hugging Face.Hugging Face model cards and repo metadata
Endpointhttps://api.ai21.com/studio/v1/chat/completionsAI21 Chat request reference code samples
ComplianceSOC 2, ISO 27001, ISO 27017, ISO 27018AI21 Jamba foundation models

The parametersEvery knob, and what moving it does

ParameterWhat it doesRange or defaultWhat happens when you move it
modelModel string.Required. jamba-large or jamba-mini, or a dated version.AI21 advises using the dated form to avoid being moved onto a new snapshot mid-project.
messagesConversation turns with system, user, assistant and tool roles.RequiredA tool message must reference an existing tool_call_id from a prior assistant turn or the request is rejected.
max_tokensOutput cap.Maximum allowed value is 4096.This is the hard ceiling on Jamba output regardless of the 256K input. Long generations must be chunked by you.
temperatureRandomness.Default 0.4. Range 0.0 to 2.0.The range goes to 2.0, wider than most enterprise APIs. Above about 1.0 output degrades quickly.
top_pNucleus sampling.Default 1.0. Range 0 to 1.0.Default of 1.0 means no truncation at all - the full tail is in play, scaled only by temperature.
stopHalt strings.Array of strings, each up to 64K characters, newlines allowedThe stop sequence is not included in the returned message.
nHow many completions to generate.Default 1. Range 1 to 16.With n above 1, temperature 0 fails outright because all answers would be duplicates. n must be 1 when streaming.
streamServer-sent event streaming.BooleanMust be False when using tools. This is the single most common Jamba integration break.
toolsFunction definitions.Array. Only type of function is supported.There are no built-in functions - every tool is yours to implement.
documentsGrounding sources with content plus key-value metadata.Array of objectsMetadata keys should be things a model understands - author, date, url - since they are surfaced to the model as text.
response_formatJSON mode.{"type": "json_object"}Guarantees valid JSON structure. AI21 documents no json_schema variant, so shape enforcement is on you.
tool_choiceNot supported. Absent from the AI21 Chat request reference.Not supportedYou cannot force a tool call. Use response_format plus an explicit instruction, or validate and retry.
top_kNot supported. Absent from the request reference.Not supportedUse top_p for truncation.
seedNot supported. Absent from the request reference.Not supportedThe nearest to determinism is temperature 0 with n of 1.
frequency_penalty and presence_penaltyNot supported. Absent from the request reference.Not supportedControl repetition with stop sequences and prompt instructions.
logprobsNot supported. Absent from the request reference.Not supportedIf you need token probabilities, self-host the open weights and read them from your serving stack.

Two knobs and one absence dominate here. max_tokens at a 4096 ceiling is the thing that reshapes designs - a 256K input with a 4K output means Jamba is a reader and a summariser, not a long-form writer, and any "rewrite this whole document" workflow has to be chunked by hand. stream being incompatible with tools is the second: agent loops that stream fine in development break the moment you attach a tool. And the missing tool_choice means you cannot force a call, so a Jamba agent needs a validation-and-retry layer that an OpenAI-shaped agent would not.

SamplingShaping the output distribution

Jamba's defaults sit at the opposite end from Cohere's. Temperature defaults to 0.4 on a range that runs all the way to 2.0, and top_p defaults to 1.0 - meaning no nucleus truncation whatsoever out of the box. The full tail of the distribution is live, moderated only by a moderately cool temperature. The practical consequence is that Jamba's failure mode under a bad prompt is an occasional very unlikely token rather than the flatness you get from an aggressively truncated default, and the fix is to set top_p yourself rather than to lower temperature further. There is no top_k to reach for. For extraction, classification and grounded QA, start at temperature 0 with top_p 1.0 - and note that at temperature 0 you must leave n at 1, since AI21 rejects the combination. For conversational use, AI21's own Jamba2 Mini model card example calls generate with temperature=0.6, which is a reasonable anchor. If output starts drifting off-vocabulary, drop top_p to 0.9 before touching temperature.

ReasoningThinking, effort and budgets

Neither jamba-large nor jamba-mini has a thinking mode. There is no thinking, reasoning_effort or budget parameter in the AI21 Chat request reference, and no reasoning content block in the response. AI21 positions this as deliberate: the Jamba2 Mini model card says the model "delivers precise question answering without the computational overhead of reasoning models". So there are no reasoning tokens to be billed for and none to display. What you do instead: prompt for explicit step-by-step working inside the visible answer, keeping in mind the 4096-token output ceiling covers the working and the answer; or move up the stack to AI21 Maestro, which AI21 documents as a planning and validated-output layer above the models rather than a mode inside them. AI21 also publishes a separate small reasoning model, ai21labs/AI21-Jamba-Reasoning-3B, on Hugging Face - its card was not reviewed for this page.

ToolsFunction calling and server tools

Function calling is available on all Jamba models through the Chat Completions endpoint. You pass JSON-schema function definitions in tools, the model returns tool_calls on the assistant message, you execute and reply with a tool role message carrying a matching tool_call_id. AI21 is explicit that only function type tools exist and that there are no built-in functions - no web search, no code interpreter. Structured output is response_format with {"type": "json_object"}, which guarantees valid JSON but not a particular schema, since AI21 documents no json_schema variant. Three failure modes bite people. First and worst: stream must be False when tools is present, so streaming agents break on tool attach. Second, there is no tool_choice, so you cannot force or forbid a call - build retries. Third, if you self-host, AI21's own vLLM command line is worth copying verbatim, including --enable-auto-tool-choice and --tool-call-parser hermes; tool calls silently fail to parse without them.

CostPrice, caching, batching, what drives the bill

Prices are from AI21's pricing page. Jamba Mini at $0.2 per million input is among the cheapest capable long-context models available, and Jamba Large at $2 in / $8 out sits roughly where Command A does. AI21 makes a specific tokenizer claim on that page - that its tokenization fits "up to 30% more text per token than other providers" - which if true means the effective price per page of text is lower than the per-token comparison suggests. The claim is AI21's own and is not independently verified here. There is no published prompt-caching discount. The Batch API exists but is enterprise-gated: AI21 says it is "currently available for enterprise use" through a sales conversation, and publishes no batch discount rate. Self-hosting shifts the whole bill to compute; AI21's own vLLM guidance suggests --quantization experts_int8 and --enable-prefix-caching, both of which matter more to your bill than any API knob.

ModelInput per 1MOutput per 1M
Jamba Mini$0.2$0.4
Jamba Large$2$8
Batch APIEnterprise only, no published rate
Prompt cachingNo published discount

Where it runsSurfaces and availability

SurfaceAvailableNotes
AI21 SaaS APIYesManaged, version 2. https://api.ai21.com/studio/v1/chat/completions. New accounts get a $10 credit good for three months.
Hugging Face weightsYesSelf-deploy, version 2. Jamba2 Mini and Jamba2 3B are ungated under Apache 2.0; Jamba Large 1.7 is gated under the Jamba Open Model License.
Hugging Face Inference ProvidersNoThe Hub API returns an empty provider mapping for AI21-Jamba2-Mini - weights only, no serverless endpoint.
KaggleYesSelf-deploy, version 1.7.
Google Vertex AI Model GardenYesSelf-deploy only, version 1.6. Jamba Mini 1.6 listed as coming soon.
Microsoft AzureYesSelf-deploy, version 1.5.
Amazon SageMakerYesSelf-deploy, version 1.5. Billed by instance-hour, not per token.
Amazon BedrockUnverifiedAI21 lists Bedrock as managed at version 1.5, but no AI21 or Jamba SKU appears in the AWS Price List for us-east-1 pulled 2026-07-23. Check the Bedrock console before planning around it.
Self-hosting via vLLMYesRequires vLLM 0.12.0 or higher, with --mamba-ssm-cache-dtype float32.

StrengthsWhat it is good at

  • The hybrid design is real and visible in config.json: one attention layer per eight, Mamba state-space blocks elsewhere, so the KV cache does not grow linearly with a 256K context the way a pure transformer's does.
  • Sparse compute where it counts - num_experts 16 with num_experts_per_tok 2 on alternating layers gives 12B active from 52B total on Jamba2 Mini.
  • Jamba2 Mini's weights are Apache 2.0, unusually permissive for a 52B model, with no acceptable-use rider and no non-commercial clause.
  • Genuinely cheap at the Mini tier - $0.2 per million input, with an AI21 claim of up to 30% more text per token than other providers.
  • Private deployment is the documented first-class path - VPC, on-premises or hybrid - backed by SOC 2 and ISO 27001, 27017 and 27018.

LimitsWhere it falls down

  • The API caps output at 4096 tokens. A 256K-in, 4K-out model is a reader, not a writer, and long-form generation has to be orchestrated by you.
  • stream must be false whenever tools is set, so you cannot stream a tool-using agent through the AI21 API.
  • No tool_choice, no seed, no top_k, no penalties and no logprobs. The parameter surface is the smallest of the four models on this page.
  • AI21's own documentation contradicts itself on recency - the Model Details block says the knowledge cutoff is August 22nd, 2024 while the Limitations block on the same page says the training dataset was created in March 2024.
  • Version drift across surfaces is severe: AI21 SaaS and Hugging Face are on version 2, Kaggle on 1.7, Vertex Model Garden on 1.6, Azure and SageMaker on 1.5. "Jamba" means a different model depending on where you run it.

Against its neighboursHow it compares

Against Falcon 3, the other open-weight efficient model here, the split is architectural: Falcon 3 is a conventional dense Llama-shaped transformer with a 32K window that runs anywhere, while Jamba trades ecosystem convenience for an 8x longer context and a memory profile that does not blow up at length - at the cost of needing vLLM 0.12+ and Mamba kernels. Against Cohere's Command A, both chase enterprise RAG in private deployments, but Command A gives you citation spans as an API feature while Jamba gives you a 4096-token output cap and a much lower price. Against Gemini 3 Flash, Jamba's argument is entirely about where the weights sit: if the data can leave your network, a hosted frontier-family small model will usually win on quality; if it cannot, Jamba2 Mini under Apache 2.0 is one of the few 50B-class models you can simply take.

Getting startedThe smallest call that works

code
curl https://api.ai21.com/studio/v1/chat/completions \
     --header "Authorization: Bearer $AI21_API_KEY" \
     --header "Content-Type: application/json" \
     --data '{
       "model": "jamba-mini",
       "max_tokens": 1024,
       "temperature": 0,
       "messages": [
         {"role": "user", "content": "Summarise the risk section."}
       ]
     }'

Pin the version first - change jamba-mini to jamba-mini-2-2026-01 so a snapshot rollover cannot change your outputs. Then raise max_tokens only up to 4096, which is the hard ceiling. If you add tools, remember to leave stream false, and add documents rather than pasting sources into the user message.

SourcesWhere every claim above came from

  • Jamba - AI21 Docs - Model Details table with Jamba Large at 398B/94B active, Jamba Mini at 52B/12B active and Jamba 3B, all at 256K; API versioning and what jamba-large and jamba-mini currently point to; deprecation dates; the August 22nd 2024 knowledge cutoff; the 9 supported languages; the Limitations section stating a March 2024 dataset; SOC 2 and ISO certifications.
  • Chat request - AI21 API reference - every request parameter: the 4096 max_tokens ceiling, temperature default 0.4 with range 0.0-2.0, top_p default 1.0, stop, n range 1-16 with its temperature and streaming constraints, stream incompatibility with tools, documents, response_format, and the endpoint URL in the code samples.
  • AI21-Jamba2-Mini config.json - hidden_size 4096, num_hidden_layers 32, num_attention_heads 32, num_key_value_heads 8, num_experts 16, num_experts_per_tok 2, attn_layer_period 8 / offset 4, expert_layer_period 2 / offset 1, mamba_d_state 16, mamba_d_conv 4, mamba_expand 2, mamba_dt_rank 256, vocab_size 65536, max_position_embeddings 262144, and the absence of any rope_theta key.
  • AI21-Jamba2-3B config.json - hidden_size 2560, 28 layers, 20 attention heads, 1 KV head, num_experts 1, attn_layer_period 14.
  • ai21labs/AI21-Jamba2-Mini model card - Apache 2.0 licence, 12B active from 52B total, the training pipeline from Jamba 1.5 through 500B mid-training tokens, state passing, cold-start SFT, DPO and on-policy RL, and the vLLM serving flags. Parameter total of 51,570,323,328 read from the repo's safetensors index.
  • Model Availability by Platform - AI21 Docs - the per-platform version matrix for AI21 SaaS, Hugging Face, Kaggle, GCP Model Garden, Azure, SageMaker and Bedrock.
  • AI21 Pricing - Jamba Mini at $0.2 / $0.4 per 1M and Jamba Large at $2 / $8 per 1M, plus the "up to 30% more text per token" tokenizer claim.
  • Pricing - AI21 Docs - per-token billing mechanics, the $10 three-month trial credit, and the note that cloud-hosted usage is billed by the cloud provider.
  • Could not confirm: Jamba Large 1.7's config.json, since the repo is gated on Hugging Face - all architecture numbers above come from the ungated Jamba2 Mini and Jamba2 3B repos; any batch or caching discount; current Bedrock availability, where AI21 claims version 1.5 but no AI21 SKU appears in the AWS us-east-1 price list. AI21's own documentation contradicts itself on the knowledge cutoff and both figures are reported above.
A living map of modern AI — kept current every morning