AI21's hybrid Mamba-Transformer family: 256K context, mixture-of-experts, Apache 2.0 weights for Jamba2 Mini, and a hard 4096-token output cap on the API.
Why this oneWhat it is actually for
You reach for Jamba when the input is long, the budget is tight and the data cannot leave your network. The hybrid architecture puts a Mamba state-space block in most layers and full attention in only one layer of every eight, so the KV cache that normally makes 256K contexts expensive largely is not there - the state is fixed size. That is the whole pitch: long-context work at a cost curve that does not blow up with sequence length. AI21 sells this primarily as a private deployment - VPC, on-premises or hybrid - with the weights on Hugging Face, Jamba2 Mini under Apache 2.0. If you are doing document QA, classification or grounded retrieval over big corpora inside a regulated network, this is the shape that fits.
What it isArchitecture, lineage, training
Jamba is a joint attention and Mamba architecture - AI21's documentation gives the model type literally as "Joint Attention and Mamba (Jamba)". Most layers are Mamba state-space blocks; attention appears periodically. Mixture-of-experts is layered on top of that, with experts on alternating layers.
The real numbers are in config.json. For ai21labs/AI21-Jamba2-Mini: hidden_size 4096, num_hidden_layers 32, num_attention_heads 32, num_key_value_heads 8, vocab_size 65536, max_position_embeddings 262144. The interleaving is set by attn_layer_period 8 with attn_layer_offset 4 - one attention layer in every eight - and expert_layer_period 2 with expert_layer_offset 1, so every other layer is a mixture-of-experts layer. Those MoE layers carry num_experts 16 with num_experts_per_tok 2. The Mamba blocks are configured with mamba_d_state 16, mamba_d_conv 4, mamba_expand 2 and mamba_dt_rank 256. There is no rope_theta key at all, which is the tell that attention here does not use rotary position embeddings - position information comes from the recurrent blocks.
The Hugging Face safetensors index reports 51,570,323,328 parameters in BF16 for Jamba2 Mini, matching AI21's "52B parameters, 12B active" claim. Jamba Large is stated at 398B total and 94B active. ai21labs/AI21-Jamba2-3B, by contrast, is dense - its config.json sets num_experts to 1 - with hidden_size 2560, 28 layers, 20 attention heads, a single KV head and attn_layer_period 14.
AI21 describes Jamba2 training as starting from Jamba 1.5 pre-training, then mid-training on 500B curated tokens weighted toward maths, code, high-quality web text and long documents, a "state passing" phase that tunes the Mamba layers for context-length generalisation, cold-start supervised fine-tuning, DPO, then several rounds of on-policy reinforcement learning moving from short-context verifiable rewards to long-context mixed rewards.
At a glanceSee it
Mamba blocks carry the long context in fixed state while sparse experts and rare attention layers do the work.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| API model strings | jamba-large currently points to jamba-large-1.7-2025-07; jamba-mini currently points to jamba-mini-2-2026-01 | AI21 Jamba foundation models page, API Versioning |
| Context window | 256K. max_position_embeddings is 262144 in config.json. | AI21 docs; AI21-Jamba2-Mini config.json |
| Max output tokens | 4096. "For Jamba models, the maximum allowed value is 4096 tokens." | AI21 Chat request reference, max_tokens |
| Model sizes | Jamba Large 398B total / 94B active; Jamba Mini 52B total / 12B active; Jamba 3B dense, no API endpoint | AI21 Jamba foundation models, Model Details table |
| Measured parameter count | 51,570,323,328 BF16 for AI21-Jamba2-Mini | Hugging Face safetensors index |
| Modalities in / out | Text in, text out | AI21 Jamba foundation models, Model Details |
| Knowledge cutoff | August 22nd, 2024 per the Model Details block. The Limitations block on the same page says the model "was trained on a dataset created in March 2024" - AI21's own page disagrees with itself. | AI21 Jamba foundation models |
| Languages | 9 officially: English, Spanish, French, Portuguese, Italian, Dutch, German, Arabic, Hebrew | AI21 Jamba foundation models |
| Weights available | Yes. ai21labs/AI21-Jamba2-Mini under Apache 2.0. ai21labs/AI21-Jamba-Large-1.7 under the Jamba Open Model License and gated on Hugging Face. | Hugging Face model cards and repo metadata |
| Endpoint | https://api.ai21.com/studio/v1/chat/completions | AI21 Chat request reference code samples |
| Compliance | SOC 2, ISO 27001, ISO 27017, ISO 27018 | AI21 Jamba foundation models |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
model | Model string. | Required. jamba-large or jamba-mini, or a dated version. | AI21 advises using the dated form to avoid being moved onto a new snapshot mid-project. |
messages | Conversation turns with system, user, assistant and tool roles. | Required | A tool message must reference an existing tool_call_id from a prior assistant turn or the request is rejected. |
max_tokens | Output cap. | Maximum allowed value is 4096. | This is the hard ceiling on Jamba output regardless of the 256K input. Long generations must be chunked by you. |
temperature | Randomness. | Default 0.4. Range 0.0 to 2.0. | The range goes to 2.0, wider than most enterprise APIs. Above about 1.0 output degrades quickly. |
top_p | Nucleus sampling. | Default 1.0. Range 0 to 1.0. | Default of 1.0 means no truncation at all - the full tail is in play, scaled only by temperature. |
stop | Halt strings. | Array of strings, each up to 64K characters, newlines allowed | The stop sequence is not included in the returned message. |
n | How many completions to generate. | Default 1. Range 1 to 16. | With n above 1, temperature 0 fails outright because all answers would be duplicates. n must be 1 when streaming. |
stream | Server-sent event streaming. | Boolean | Must be False when using tools. This is the single most common Jamba integration break. |
tools | Function definitions. | Array. Only type of function is supported. | There are no built-in functions - every tool is yours to implement. |
documents | Grounding sources with content plus key-value metadata. | Array of objects | Metadata keys should be things a model understands - author, date, url - since they are surfaced to the model as text. |
response_format | JSON mode. | {"type": "json_object"} | Guarantees valid JSON structure. AI21 documents no json_schema variant, so shape enforcement is on you. |
tool_choice | Not supported. Absent from the AI21 Chat request reference. | Not supported | You cannot force a tool call. Use response_format plus an explicit instruction, or validate and retry. |
top_k | Not supported. Absent from the request reference. | Not supported | Use top_p for truncation. |
seed | Not supported. Absent from the request reference. | Not supported | The nearest to determinism is temperature 0 with n of 1. |
frequency_penalty and presence_penalty | Not supported. Absent from the request reference. | Not supported | Control repetition with stop sequences and prompt instructions. |
logprobs | Not supported. Absent from the request reference. | Not supported | If you need token probabilities, self-host the open weights and read them from your serving stack. |
Two knobs and one absence dominate here. max_tokens at a 4096 ceiling is the thing that reshapes designs - a 256K input with a 4K output means Jamba is a reader and a summariser, not a long-form writer, and any "rewrite this whole document" workflow has to be chunked by hand. stream being incompatible with tools is the second: agent loops that stream fine in development break the moment you attach a tool. And the missing tool_choice means you cannot force a call, so a Jamba agent needs a validation-and-retry layer that an OpenAI-shaped agent would not.
SamplingShaping the output distribution
Jamba's defaults sit at the opposite end from Cohere's. Temperature defaults to 0.4 on a range that runs all the way to 2.0, and top_p defaults to 1.0 - meaning no nucleus truncation whatsoever out of the box. The full tail of the distribution is live, moderated only by a moderately cool temperature. The practical consequence is that Jamba's failure mode under a bad prompt is an occasional very unlikely token rather than the flatness you get from an aggressively truncated default, and the fix is to set top_p yourself rather than to lower temperature further. There is no top_k to reach for. For extraction, classification and grounded QA, start at temperature 0 with top_p 1.0 - and note that at temperature 0 you must leave n at 1, since AI21 rejects the combination. For conversational use, AI21's own Jamba2 Mini model card example calls generate with temperature=0.6, which is a reasonable anchor. If output starts drifting off-vocabulary, drop top_p to 0.9 before touching temperature.
ReasoningThinking, effort and budgets
Neither jamba-large nor jamba-mini has a thinking mode. There is no thinking, reasoning_effort or budget parameter in the AI21 Chat request reference, and no reasoning content block in the response. AI21 positions this as deliberate: the Jamba2 Mini model card says the model "delivers precise question answering without the computational overhead of reasoning models". So there are no reasoning tokens to be billed for and none to display. What you do instead: prompt for explicit step-by-step working inside the visible answer, keeping in mind the 4096-token output ceiling covers the working and the answer; or move up the stack to AI21 Maestro, which AI21 documents as a planning and validated-output layer above the models rather than a mode inside them. AI21 also publishes a separate small reasoning model, ai21labs/AI21-Jamba-Reasoning-3B, on Hugging Face - its card was not reviewed for this page.
ToolsFunction calling and server tools
Function calling is available on all Jamba models through the Chat Completions endpoint. You pass JSON-schema function definitions in tools, the model returns tool_calls on the assistant message, you execute and reply with a tool role message carrying a matching tool_call_id. AI21 is explicit that only function type tools exist and that there are no built-in functions - no web search, no code interpreter. Structured output is response_format with {"type": "json_object"}, which guarantees valid JSON but not a particular schema, since AI21 documents no json_schema variant. Three failure modes bite people. First and worst: stream must be False when tools is present, so streaming agents break on tool attach. Second, there is no tool_choice, so you cannot force or forbid a call - build retries. Third, if you self-host, AI21's own vLLM command line is worth copying verbatim, including --enable-auto-tool-choice and --tool-call-parser hermes; tool calls silently fail to parse without them.
CostPrice, caching, batching, what drives the bill
Prices are from AI21's pricing page. Jamba Mini at $0.2 per million input is among the cheapest capable long-context models available, and Jamba Large at $2 in / $8 out sits roughly where Command A does. AI21 makes a specific tokenizer claim on that page - that its tokenization fits "up to 30% more text per token than other providers" - which if true means the effective price per page of text is lower than the per-token comparison suggests. The claim is AI21's own and is not independently verified here. There is no published prompt-caching discount. The Batch API exists but is enterprise-gated: AI21 says it is "currently available for enterprise use" through a sales conversation, and publishes no batch discount rate. Self-hosting shifts the whole bill to compute; AI21's own vLLM guidance suggests --quantization experts_int8 and --enable-prefix-caching, both of which matter more to your bill than any API knob.
| Model | Input per 1M | Output per 1M |
|---|---|---|
| Jamba Mini | $0.2 | $0.4 |
| Jamba Large | $2 | $8 |
| Batch API | Enterprise only, no published rate | |
| Prompt caching | No published discount | |
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| AI21 SaaS API | Yes | Managed, version 2. https://api.ai21.com/studio/v1/chat/completions. New accounts get a $10 credit good for three months. |
| Hugging Face weights | Yes | Self-deploy, version 2. Jamba2 Mini and Jamba2 3B are ungated under Apache 2.0; Jamba Large 1.7 is gated under the Jamba Open Model License. |
| Hugging Face Inference Providers | No | The Hub API returns an empty provider mapping for AI21-Jamba2-Mini - weights only, no serverless endpoint. |
| Kaggle | Yes | Self-deploy, version 1.7. |
| Google Vertex AI Model Garden | Yes | Self-deploy only, version 1.6. Jamba Mini 1.6 listed as coming soon. |
| Microsoft Azure | Yes | Self-deploy, version 1.5. |
| Amazon SageMaker | Yes | Self-deploy, version 1.5. Billed by instance-hour, not per token. |
| Amazon Bedrock | Unverified | AI21 lists Bedrock as managed at version 1.5, but no AI21 or Jamba SKU appears in the AWS Price List for us-east-1 pulled 2026-07-23. Check the Bedrock console before planning around it. |
| Self-hosting via vLLM | Yes | Requires vLLM 0.12.0 or higher, with --mamba-ssm-cache-dtype float32. |
StrengthsWhat it is good at
- The hybrid design is real and visible in
config.json: one attention layer per eight, Mamba state-space blocks elsewhere, so the KV cache does not grow linearly with a 256K context the way a pure transformer's does. - Sparse compute where it counts -
num_experts16 withnum_experts_per_tok2 on alternating layers gives 12B active from 52B total on Jamba2 Mini. - Jamba2 Mini's weights are Apache 2.0, unusually permissive for a 52B model, with no acceptable-use rider and no non-commercial clause.
- Genuinely cheap at the Mini tier - $0.2 per million input, with an AI21 claim of up to 30% more text per token than other providers.
- Private deployment is the documented first-class path - VPC, on-premises or hybrid - backed by SOC 2 and ISO 27001, 27017 and 27018.
LimitsWhere it falls down
- The API caps output at 4096 tokens. A 256K-in, 4K-out model is a reader, not a writer, and long-form generation has to be orchestrated by you.
streammust be false whenevertoolsis set, so you cannot stream a tool-using agent through the AI21 API.- No
tool_choice, noseed, notop_k, no penalties and nologprobs. The parameter surface is the smallest of the four models on this page. - AI21's own documentation contradicts itself on recency - the Model Details block says the knowledge cutoff is August 22nd, 2024 while the Limitations block on the same page says the training dataset was created in March 2024.
- Version drift across surfaces is severe: AI21 SaaS and Hugging Face are on version 2, Kaggle on 1.7, Vertex Model Garden on 1.6, Azure and SageMaker on 1.5. "Jamba" means a different model depending on where you run it.
Against its neighboursHow it compares
Against Falcon 3, the other open-weight efficient model here, the split is architectural: Falcon 3 is a conventional dense Llama-shaped transformer with a 32K window that runs anywhere, while Jamba trades ecosystem convenience for an 8x longer context and a memory profile that does not blow up at length - at the cost of needing vLLM 0.12+ and Mamba kernels. Against Cohere's Command A, both chase enterprise RAG in private deployments, but Command A gives you citation spans as an API feature while Jamba gives you a 4096-token output cap and a much lower price. Against Gemini 3 Flash, Jamba's argument is entirely about where the weights sit: if the data can leave your network, a hosted frontier-family small model will usually win on quality; if it cannot, Jamba2 Mini under Apache 2.0 is one of the few 50B-class models you can simply take.
Getting startedThe smallest call that works
curl https://api.ai21.com/studio/v1/chat/completions \
--header "Authorization: Bearer $AI21_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"model": "jamba-mini",
"max_tokens": 1024,
"temperature": 0,
"messages": [
{"role": "user", "content": "Summarise the risk section."}
]
}'Pin the version first - change jamba-mini to jamba-mini-2-2026-01 so a snapshot rollover cannot change your outputs. Then raise max_tokens only up to 4096, which is the hard ceiling. If you add tools, remember to leave stream false, and add documents rather than pasting sources into the user message.
SourcesWhere every claim above came from
- Jamba - AI21 Docs - Model Details table with Jamba Large at 398B/94B active, Jamba Mini at 52B/12B active and Jamba 3B, all at 256K; API versioning and what
jamba-largeandjamba-minicurrently point to; deprecation dates; the August 22nd 2024 knowledge cutoff; the 9 supported languages; the Limitations section stating a March 2024 dataset; SOC 2 and ISO certifications. - Chat request - AI21 API reference - every request parameter: the 4096
max_tokensceiling,temperaturedefault 0.4 with range 0.0-2.0,top_pdefault 1.0,stop,nrange 1-16 with its temperature and streaming constraints,streamincompatibility withtools,documents,response_format, and the endpoint URL in the code samples. - AI21-Jamba2-Mini config.json - hidden_size 4096, num_hidden_layers 32, num_attention_heads 32, num_key_value_heads 8, num_experts 16, num_experts_per_tok 2, attn_layer_period 8 / offset 4, expert_layer_period 2 / offset 1, mamba_d_state 16, mamba_d_conv 4, mamba_expand 2, mamba_dt_rank 256, vocab_size 65536, max_position_embeddings 262144, and the absence of any rope_theta key.
- AI21-Jamba2-3B config.json - hidden_size 2560, 28 layers, 20 attention heads, 1 KV head, num_experts 1, attn_layer_period 14.
- ai21labs/AI21-Jamba2-Mini model card - Apache 2.0 licence, 12B active from 52B total, the training pipeline from Jamba 1.5 through 500B mid-training tokens, state passing, cold-start SFT, DPO and on-policy RL, and the vLLM serving flags. Parameter total of 51,570,323,328 read from the repo's safetensors index.
- Model Availability by Platform - AI21 Docs - the per-platform version matrix for AI21 SaaS, Hugging Face, Kaggle, GCP Model Garden, Azure, SageMaker and Bedrock.
- AI21 Pricing - Jamba Mini at $0.2 / $0.4 per 1M and Jamba Large at $2 / $8 per 1M, plus the "up to 30% more text per token" tokenizer claim.
- Pricing - AI21 Docs - per-token billing mechanics, the $10 three-month trial credit, and the note that cloud-hosted usage is billed by the cloud provider.
- Could not confirm: Jamba Large 1.7's config.json, since the repo is gated on Hugging Face - all architecture numbers above come from the ungated Jamba2 Mini and Jamba2 3B repos; any batch or caching discount; current Bedrock availability, where AI21 claims version 1.5 but no AI21 SKU appears in the AWS us-east-1 price list. AI21's own documentation contradicts itself on the knowledge cutoff and both figures are reported above.