Home › Frontier Models › Falcon 3
Model · Reference

Falcon 3

Fully-open self-host

In one line

TII's December 2024 open-weight family, 1B to 10B, 32K context, no hosted API - you run it yourself under the TII Falcon licence, which is not Apache.

Why this oneWhat it is actually for

You reach for Falcon 3 when the model has to run entirely on hardware you control and nothing may leave the building. It is small enough that the 10B fits in about 20 GB of weights in bf16 - one 24 GB card - and its config.json declares LlamaForCausalLM, so it drops into transformers, vLLM, llama.cpp and every Llama-shaped serving path with no custom code. The instruct models ship a chat template with a built-in function-calling branch, and TII post-trained on function-call data, so tool use is not bolted on afterwards. Its honest position: this is a late-2024 model in a 2026 field. Pick it for sovereignty, air-gapped deployment, edge inference or fine-tuning substrate - not for reasoning quality.

What it isArchitecture, lineage, training

Falcon 3 is a dense, decoder-only transformer. There is no mixture-of-experts anywhere in the line. TII published four sizes - 1B, 3B, 7B and 10B - each as a base and an instruct model, plus a Falcon3-Mamba-7B variant on a different architecture.

Numbers below are from tiiuae/Falcon3-10B-Instruct/config.json unless stated. architectures is LlamaForCausalLM. hidden_size 3072, num_hidden_layers 40, num_attention_heads 12, num_key_value_heads 4 - grouped-query attention at a 3-to-1 ratio - and an unusually wide head_dim of 256, so the attention dimension is 3072 rather than matching hidden size. intermediate_size is 23040 with SwiGLU activation and RMSNorm, vocab_size is 131072, rope_theta is 1000042 and max_position_embeddings is 32768 with rope_scaling null - the 32K window is native, not extended. The safetensors index gives 10,305,653,760 parameters in BF16.

Falcon3-7B-Instruct shares every width with the 10B and differs only in depth at 28 layers, because the 10B was depth up-scaled from Falcon3-7B-Base - TII duplicated and continued training layers rather than training a wider model. The card states that up-scaling consumed 2 teratokens of web, code, STEM, high-quality and multilingual data on 1024 H100 chips, and that post-training used 1.2 million samples spanning STEM, conversational, code, safety and function-call data. TII's own Falcon 3 page puts family-wide pretraining at 14 trillion tokens.

Falcon3-1B-Instruct is a different shape - hidden_size 2048, 18 layers, 8 heads over 4 KV heads, intermediate_size 8192 and max_position_embeddings of only 8192. The 32K context claim applies to the larger members, not the whole family.

Languages are four: English, French, Spanish, Portuguese. Release date on the card is December 2024. Note the site row said 2026 - that is wrong.

At a glanceSee it

Falcon 3 diagram

Falcon 3 is a Llama-shaped dense model you serve yourself; TII ships weights, not an endpoint.

CapacityContext, output and what fits

FactValueSource
Context window32,768 tokens. max_position_embeddings 32768, rope_scaling null, tokenizer model_max_length 32768. The 1B variant is 8192.Falcon3-10B-Instruct config.json and tokenizer_config.json; Falcon3-1B-Instruct config.json
Max output tokensNo separate cap. Output is bounded by the 32K window minus your prompt, and by whatever max_new_tokens you set.No fixed limit stated on the model card
Parameters10,305,653,760 BF16 for Falcon3-10B-InstructHugging Face safetensors index
ArchitectureDense decoder-only, declared as LlamaForCausalLM. hidden_size 3072, 40 layers, 12 query heads, 4 KV heads, head_dim 256, intermediate_size 23040, vocab_size 131072, rope_theta 1000042, SwiGLU and RMSNorm.config.json; model card Model Details
Modalities in / outText in, text outModel card
Languages4: English, French, Spanish, PortugueseModel card; TII Falcon 3 page
Training data10B depth up-scaled from Falcon3-7B-Base with 2 teratokens on 1024 H100 chips; post-trained on 1.2M samples. TII states 14 trillion tokens across the family.Model card; TII Falcon 3 page
Knowledge cutoffNot disclosedNeither the model card nor the TII page states one
Release dateDecember 2024Model card, Model Release Date
Weights availableYes, ungated. Licence is falcon-llm-license, named on the card as TII Falcon-LLM License 2.0. Based in part on Apache 2.0 with modifications; commercial use permitted, but with a mandatory Acceptable Use Policy and a mandatory attribution statement for derivative works. Not an OSI-approved licence - read the terms.Model card metadata; TII Falcon terms and conditions
Special tokenseos_token is <|endoftext|>, pad_token is <|pad|>, bos_token is nulltokenizer_config.json
Benchmarks reported on the cardIFEval 78.17, BBH 3-shot 44.82, MATH Lvl-5 4-shot 25.91, GPQA 0-shot 10.51, MuSR 0-shot 13.61, MMLU-PRO 5-shot 38.1Model card model-index, self-reported against the Open LLM Leaderboard

The parametersEvery knob, and what moving it does

There is no TII inference API, so there are no vendor request parameters. The knobs below are the Hugging Face transformers GenerationConfig fields you pass to model.generate(), plus the one argument Falcon 3's own chat template consumes. If you serve with vLLM or llama.cpp instead, the names are that server's; those were not verified for this page.

ParameterWhat it doesRange or defaultWhat happens when you move it
max_new_tokensTokens to generate, ignoring prompt length.Optional, no default. Recommended over max_length.Unset, generation runs to the model's default length. Always set it.
do_sampleSampling versus greedy decoding.Boolean. Falcon 3 ships no value, so the library default of greedy applies.This is the important one. Falcon3-10B-Instruct's generation_config.json contains only bos and eos ids - no sampling defaults at all - so out of the box it decodes greedily and looks deterministic.
temperatureScales next-token logits.Read from generation_config.json; if unset the library default is 1.0. Falcon 3 does not set it.Has no effect at all unless do_sample=True. A very common silent no-op.
top_pNucleus sampling.Library default 1.0 when unset. Falcon 3 does not set it.0.9 to 0.95 is the usual starting band once sampling is on.
top_kTop-k truncation.Library default 50 when unset. Falcon 3 does not set it.Because the library default is already 50, Falcon 3 sampling is k-truncated whether you asked for it or not. Set top_k=0 to disable.
min_pMinimum token probability, scaled by the top token's probability.0 to 1. Typical values 0.01 to 0.2.0.05 to 0.1 is roughly as selective as top_p 0.9 and behaves better at high temperature.
repetition_penaltyPenalises already-generated tokens.1.0 means no penalty.1.05 to 1.15 fixes most looping. Above 1.2 the model starts avoiding necessary words.
no_repeat_ngram_sizeForbids repeating any n-gram of this size.Integer above 0 to enable.Blunt. It will also block legitimate repeated phrases like a recurring defined term.
num_beamsBeam search width.1 means no beam search.Above 1 with do_sample=False gives beam search; multiplies compute by the beam count.
num_return_sequencesHow many completions to return per input.IntegerWith sampling this gives you n-best for a validator to choose from.
stop_stringsStrings that terminate generation.String or list of stringsNeeds the tokenizer passed to generate to work.
eos_token_idEnd-of-sequence token.Falcon 3 sets 11, which is <|endoftext|>.Add extra ids as a list if your fine-tune introduces new turn terminators.
pad_token_idPadding token.Tokenizer defines <|pad|>; generation_config.json does not set an id.Batched generation warns or misbehaves until you set this explicitly.
tools on apply_chat_templateInjects Falcon's function-calling system preamble and your tool schemas.List of tool definitionsThis is a tokenizer argument, not a generate argument. Without it the model never sees the function-calling preamble it was trained against.
seedNot a generate parameter.Not supported on GenerationConfigUse transformers.set_seed() before the call, or your serving stack's own seed field.
max_lengthTotal sequence length including prompt.Kept for backward compatibility.Use max_new_tokens instead - max_length silently shrinks your output as the prompt grows.

Two things matter more than the rest. First, do_sample: Falcon 3's shipped generation_config.json sets nothing but token ids, so every temperature and top_p you pass is ignored until you turn sampling on - and TII's own quickstart calls generate with only max_new_tokens, which is greedy. Second, tools on apply_chat_template: the model's function-calling behaviour is baked into the chat template, so passing tool schemas any other way throws away the post-training TII paid for.

SamplingShaping the output distribution

Falcon 3's decoding behaviour is defined by an absence. generation_config.json in Falcon3-10B-Instruct contains only bos_token_id, eos_token_id and a transformers version stamp - no do_sample, no temperature, no top_p, no top_k. TII made no recommendation, so what you get is whatever your library defaults to. In transformers that means greedy decoding, and any sampling arguments you pass are silently ignored until do_sample=True. Once you turn sampling on, the library-level fallbacks apply: temperature 1.0, top_p 1.0, top_k 50. That last one is worth noticing - top-k truncation at 50 is on by default in transformers even though nothing in the model asked for it. A reasonable starting point for a 10B instruct model of this generation is do_sample=True with temperature=0.7, top_p=0.9, top_k=0 to disable the implicit k-cut, and repetition_penalty=1.05. For extraction and classification, leave it greedy - that is the default anyway, and it is the closest thing to determinism you get without also fixing a seed.

ReasoningThinking, effort and budgets

Falcon 3 has no thinking or reasoning mode. There is no reasoning-token concept, no budget, nothing to bill and nothing to hide or display. The model card describes strong results on reasoning benchmarks at release time - IFEval 78.17, BBH 44.82, MMLU-PRO 38.1 as self-reported against the Open LLM Leaderboard - but those are ordinary single-pass generations, not extended thinking. What you do instead is prompt for it: ask for numbered working before the answer, or run your own scaffold that samples several completions and votes. Both cost real tokens against a 32K window, which is the practical ceiling on how much scratch work you can afford. If you want a Falcon-family model built for reasoning, TII's Hugging Face organisation lists a separate tiiuae/Falcon-H1R-7B repository alongside the broader Falcon-H1 line that has superseded Falcon 3 for most sizes; neither card was reviewed for this page.

ToolsFunction calling and server tools

Function calling is trained in and delivered through the chat template rather than an API. tokenizer_config.json contains a Jinja template with an explicit tools branch that emits a Falcon-specific system preamble - "You are a Falcon assistant skilled in function calling" - followed by your function schemas, and TII's post-training set of 1.2 million samples explicitly included function-call data. So the correct way to use tools is tokenizer.apply_chat_template(messages, tools=[...], add_generation_prompt=True); passing schemas any other way skips the format the model was trained on. Everything after that is yours. There is no server enforcing the schema, no tool_choice, no parallel-call protocol, and no JSON mode unless your serving stack supplies constrained decoding. Two practical failure modes: the tokenizer defines no bos_token, so code that assumes one and prepends it corrupts the prompt; and generation_config.json sets no pad_token_id even though the tokenizer defines <|pad|>, which makes batched generation warn or produce garbage until you set it yourself.

CostPrice, caching, batching, what drives the bill

There is no list price. TII publishes no hosted inference API, so the only costs are hardware and operations. Two figures set your floor, both computed from the model's own files. Weights: 10,305,653,760 parameters at 2 bytes each in BF16 is about 20.6 GB, so a single 24 GB card holds the model but leaves little room - INT8 or 4-bit quantisation, and TII publishes 1.58-bit and GPTQ variants, moves it comfortably onto consumer hardware. KV cache: grouped-query attention with 4 KV heads at head_dim 256 over 40 layers means 2 x 40 x 4 x 256 x 2 bytes = 160 KiB per token, so a full 32K context costs about 5 GiB of cache per concurrent sequence. That second number is the one that decides your batch size, and it is the main reason GQA is here at all - without it, at 12 KV heads, the same cache would be three times larger.

Cost driverFigureBasis
List price per tokenNone - no vendor APITII publishes no inference endpoint
Weights in BF16~20.6 GB10,305,653,760 params x 2 bytes, from the safetensors index
KV cache per token~160 KiBComputed from config.json: 2 x 40 layers x 4 KV heads x 256 head_dim x 2 bytes
KV cache at full 32K context~5 GiB per sequenceSame calculation x 32768
Licence costZero, with obligationsTII Falcon-LLM License 2.0 - attribution statement and Acceptable Use Policy compliance are mandatory

Where it runsSurfaces and availability

SurfaceAvailableNotes
A TII hosted APINoTII publishes weights and a demo chat at chat.falconllm.tii.ae, not a paid inference endpoint.
Hugging Face weightsYestiiuae/Falcon3-10B-Instruct and siblings at 1B, 3B and 7B, base and instruct, ungated. GGUF, GPTQ-Int8 and 1.58-bit variants are published too.
Hugging Face Inference ProvidersNoThe Hub API returns an empty provider mapping for Falcon3-10B-Instruct - no serverless endpoint is wired up.
Amazon BedrockNoNo Falcon SKU appears in the AWS Price List for AmazonBedrock us-east-1 pulled 2026-07-23.
Google Vertex AIUnverifiedCould not be confirmed from a primary source as of 2026-07-25.
Microsoft Foundry / AzureUnverifiedCould not be confirmed from a primary source as of 2026-07-25.
Self-hostingYesThe point of the model. config.json declares LlamaForCausalLM, so transformers, vLLM and llama.cpp all load it without custom code.

StrengthsWhat it is good at

  • Declares LlamaForCausalLM in config.json, so it loads in every Llama-compatible stack with no custom modelling code, no trust_remote_code and no kernel prerequisites.
  • Native 32K context with rope_scaling null and rope_theta 1000042 - the window was trained, not stretched at inference time.
  • Grouped-query attention at 12 query heads over 4 KV heads keeps the cache at about 160 KiB per token, which is what makes a 32K context practical on one card.
  • Function calling is in the shipped chat template and in the post-training mix - 1.2 million samples covering STEM, conversational, code, safety and function-call data.
  • Four sizes from 1B to 10B with base and instruct variants, plus GGUF, GPTQ and 1.58-bit releases, so the same family covers a laptop and a server.

LimitsWhere it falls down

  • Released December 2024. In a 2026 field it is well behind on reasoning and coding, and TII's own Falcon-H1 line has largely superseded it.
  • Four languages only - English, French, Spanish, Portuguese. Anything else is out of distribution.
  • No hosted API from TII and no Hugging Face inference provider mapping, so "try it quickly" means renting a GPU.
  • The licence is TII Falcon-LLM License 2.0, not Apache 2.0. Commercial use is allowed but attribution for derivative works and Acceptable Use Policy compliance are mandatory, and the AUP can be updated by TII.
  • No knowledge cutoff is published, no reasoning mode exists, and the shipped generation_config.json gives no sampling guidance at all - you are calibrating from scratch.

Against its neighboursHow it compares

Against Jamba, the other open-weight efficient model here, the trade is context versus convenience: Jamba2 Mini gives you 256K and Apache 2.0 but needs Mamba kernels and vLLM 0.12+, while Falcon 3 gives you 32K in a stock Llama shape that runs anywhere today. Against Microsoft's Phi-5, both target the small-and-local slot; Phi's argument has always been benchmark-per-parameter from curated training data, Falcon's is a genuinely permissive-ish licence from a non-US institute plus a 131K vocabulary that handles non-English text more evenly. Against Llama 5, there is no contest on capability - Falcon 3 is a generation behind - and the reason to choose it is jurisdictional: TII is Emirati, the licence terms are its own, and for organisations that need a model not governed by US or Chinese vendor terms that is the whole argument.

Getting startedThe smallest call that works

code
from transformers import AutoTokenizer, AutoModelForCausalLM

model_name = "tiiuae/Falcon3-10B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype="auto", device_map="auto"
)

messages = [
    {"role": "system", "content": "Answer in one sentence."},
    {"role": "user", "content": "How many hours in one day?"},
]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer([text], return_tensors="pt").to(model.device)

out = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.batch_decode(out[:, inputs.input_ids.shape[1]:],
                             skip_special_tokens=True)[0])

Change the decoding first. As written this is greedy, because Falcon 3's generation_config.json sets no sampling defaults - add do_sample=True, temperature=0.7, top_p=0.9, top_k=0 if you want variety. Set pad_token_id=tokenizer.pad_token_id before you batch. For tools, pass tools=[...] into apply_chat_template so the model gets the function-calling preamble it was trained on.

SourcesWhere every claim above came from

  • tiiuae/Falcon3-10B-Instruct model card - 40 decoder blocks, GQA with 12 query and 4 KV heads, head dimension 256, RoPE value 1000042, 32K context, 131K vocab, depth up-scaling from Falcon3-7B-Base with 2 teratokens on 1024 H100 chips, 1.2M post-training samples including function-call data, four supported languages, TII Falcon-LLM License 2.0, December 2024 release, and the self-reported Open LLM Leaderboard figures.
  • Falcon3-10B-Instruct config.json - architectures LlamaForCausalLM, hidden_size 3072, num_hidden_layers 40, num_attention_heads 12, num_key_value_heads 4, head_dim 256, intermediate_size 23040, vocab_size 131072, rope_theta 1000042, rope_scaling null, max_position_embeddings 32768.
  • Falcon3-7B-Instruct config.json and Falcon3-1B-Instruct config.json - confirming the 7B is the same width at 28 layers and the 1B is hidden_size 2048, 18 layers, max_position_embeddings 8192.
  • Falcon3-10B-Instruct generation_config.json - contains only bos_token_id, eos_token_id and a version stamp, which is the basis for every claim above about greedy defaults.
  • Falcon3-10B-Instruct tokenizer_config.json - eos <|endoftext|>, pad <|pad|>, no bos, model_max_length 32768, and the chat template's function-calling branch.
  • transformers GenerationConfig reference - every parameter name, and the stated fallbacks used when a model does not set them: temperature 1.0, top_k 50, top_p 1.0, and greedy decoding when do_sample is false.
  • Falcon 3 - Technology Innovation Institute - four sizes, native 32K training, 14 trillion tokens, four languages, TII Falcon License 2.0, and a demo chat rather than a paid API.
  • Falcon terms and conditions - the licence is based in part on Apache 2.0 with modifications, permits commercial use, and makes both the Acceptable Use Policy and a derivative-work attribution statement mandatory.
  • Could not confirm: any knowledge cutoff date for Falcon 3; Vertex AI or Azure availability; the capabilities of Falcon-H1R-7B, whose repository exists in TII's Hugging Face organisation but whose card was not read. The parameter total of 10,305,653,760 and the empty inference-provider mapping come from the Hugging Face Hub model API for this repository. The site row for this model said 2026 and "permissive licence" - the card says December 2024, and the licence carries attribution and acceptable-use obligations.
A living map of modern AI — kept current every morning