There is no Phi-5. Microsoft's newest Phi is Phi-4-reasoning-vision-15B, a 15B dense MIT-licensed vision reasoner with a 16,384-token window, released 4 March 2026.
Why this oneWhat it is actually for
You reach for a small open model when you want vision and some reasoning on hardware you control, at a cost that is a GPU-hour rather than a per-token bill. Phi-5 is not that model, because it does not exist. The model that actually answers the question is Phi-4-reasoning-vision-15B: 15B parameters, MIT licence, bf16 weights that fit a single 80GB card, and a hybrid design that reasons at length on a chart or a maths diagram and answers immediately on a screenshot. Reach for it when the input is an image and the output needs a defensible chain of steps — document analysis, GUI grounding for a computer-use agent, scientific figures. Its 16,384-token window is the constraint that will decide whether it fits.
What it isArchitecture, lineage, training
Phi-4-Reasoning-Vision-15B is a dense decoder with a bolted-on vision tower — no mixture of experts anywhere in it. Microsoft calls the shape mid-fusion: a SigLIP-2 encoder turns the image into visual tokens, an MLP projector maps them into the language model's embedding space, and they are injected into a pretrained text backbone. The backbone is Phi-4-Reasoning, itself descended from phi-4.
The real numbers come from config.json in the microsoft/Phi-4-reasoning-vision-15B repository: hidden_size 5120, num_hidden_layers 40, num_attention_heads 40, num_key_value_heads 10 — grouped-query attention at a four-to-one ratio — intermediate_size 17920, vocab_size 100352, rope_theta 500000, max_position_embeddings 32768, sliding_window null, dtype bfloat16, model_type "phi4-siglip". The nested vision_config gives 27 layers at hidden size 1152 with 16 heads, and mm_vision_tower names google/siglip2-so400m-patch16-naflex. Dynamic resolution is bounded by min_num_patches 256 and max_num_patches 3600.
Put that beside microsoft/phi-4's own config.json — hidden 5120, 40 layers, 40 heads, 10 KV heads, intermediate 17920 — and the lineage is literally visible. This is the phi-4 body with rope theta raised from 250000 to 500000 and a vision stack added, which is why it lands at 15B rather than 14B.
Training was supervised fine-tuning only, on a curated mix of reasoning and non-reasoning data — no RL stage is described. Microsoft Research puts the multimodal budget at 200 billion tokens, against competitors it says use more than a trillion, and the model card records 240 NVIDIA B200 GPUs for four days, with training dates spanning 2025-02-03 to 2026-02-21.
At a glanceSee it
A SigLIP-2 tower feeds a Phi-4-Reasoning body that answers in THINK or NOTHINK mode.
CapacityContext, output and what fits
| Fact | Value | Source |
|---|---|---|
| A model named Phi-5 | Does not exist. No Phi-5 repository under microsoft on Hugging Face, and no mention on Microsoft's Phi product page | Hugging Face models API for the microsoft org and azure.microsoft.com/products/phi, both checked 2026-07-25 |
| Newest Phi language model | Phi-4-reasoning-vision-15B, released 4 March 2026 | Model card header, huggingface.co/microsoft/Phi-4-reasoning-vision-15B |
| Parameters | 15B, dense | Model card header; config.json shows no expert fields |
| Context window | 16,384 tokens per the model card. config.json sets max_position_embeddings to 32768 and tokenizer_model_max_length to 16384 — treat 16,384 as the supported figure | Model card and config.json |
| Max output tokens | Not disclosed as a separate limit. Microsoft's own sample_inference.py uses max_new_tokens=4096 | sample_inference.py in the model repository |
| Modalities in | Text and images | Model card header |
| Modalities out | Text | Model card header |
| Knowledge cutoff | Not disclosed. The card gives training dates of 3 February 2025 to 21 February 2026, which is not the same thing | Model card header |
| Weights available | Yes — seven safetensors shards, plus custom modelling code | Hugging Face repository file listing |
| Licence | MIT | Model card header and repository metadata |
| Vision encoder | SigLIP-2 so400m patch16 naflex, 27 layers, hidden 1152, 16 heads; 256 to 3,600 visual tokens | config.json, mm_vision_tower and vision_config |
| Language other than English | Not supported. The card states the model is not intended to support multilingual use | Model card, Responsible AI Considerations |
The parametersEvery knob, and what moving it does
| Parameter | What it does | Range or default | What happens when you move it |
|---|---|---|---|
temperature | Scales logits before sampling | 0.8 in the shipped generation_config.json | Below about 0.3 the THINK traces get terse and repetitive; above 1.0 grounding coordinates start drifting on GUI tasks |
top_p | Nucleus cut-off | 0.95 in generation_config.json | Tightening it toward 0.8 stabilises OCR and coordinate output at the cost of shorter reasoning |
do_sample | Whether to sample at all or take the argmax | true in generation_config.json, but Microsoft's own sample_inference.py passes False | Set false and you get greedy decoding — reproducible, and what Microsoft's example actually demonstrates |
max_new_tokens | Generation ceiling | Not set in generation_config.json; Microsoft's sample uses 4096 | THINK-mode traces are long. Set this too low and you truncate inside the reasoning block before any answer appears |
top_k | Fixed candidate pool | Not set by Microsoft; whatever your serving framework defaults to | Microsoft states no value, so any number you pick is your choice, not a vendor recommendation |
eos_token_id | Stops generation | [100265] | Override it and the model runs past its own turn boundary into template noise |
pad_token_id | Padding token for batching | 100349 | Get this wrong when batching and padded positions leak into attention |
bos_token_id | Sequence start token | 100257 | Handled by the chat template; only relevant if you build prompts by hand |
<think> generation prefix | Forces extended chain-of-thought | Append after <|im_start|>assistant<|im_sep|> | Forces reasoning even on perception tasks. Microsoft's own table shows this lowering AI2D from 84.8 to 79.7 |
<nothink> generation prefix | Forces direct inference | Append after <|im_start|>assistant<|im_sep|> | Cuts latency sharply, but costs 6.8 points on ChartQA and 6.5 on MathVista-MINI in Microsoft's table |
| System prompt | Selects automatic THINK / NOTHINK routing | The card prints the exact required prompt | The card says always use the chat template and system prompt. Omit it and mode routing is undefined behaviour |
trust_remote_code | Loads the repo's custom modelling code | True in Microsoft's sample | Required — config.json maps the model through modeling_phi4_visionr.py via auto_map. Without it the load fails |
dtype | Weight precision at load | torch.bfloat16 in Microsoft's sample; config.json declares bfloat16 | Microsoft recommends serving in bf16. Quantise below that and no published numbers apply any more |
| Image resolution budget | How many visual tokens an image consumes | 256 to 3,600 visual tokens per image. The vision tower is google/siglip2-so400m-patch16-naflex, so patches are 16x16; Phi's config.json declares no patch_size of its own. Projector is mlp2x_gelu | Small images give the model too little to ground on; large ones eat the 16,384-token budget before the text arrives |
seed | Reproducibility | Not specified by Microsoft | With do_sample=False you get determinism without needing one |
tools / tool_choice / response_format | Function calling and structured output | Not supported by the model card | The card describes no tool-calling format. Constrain output in your serving stack with a grammar, or parse the text |
Two knobs matter more than the rest. The first is not a sampling parameter at all — it is whether you let the system prompt route THINK and NOTHINK automatically or force a mode with a prefix token. Microsoft's own benchmark tables show automatic routing beating both forced modes on most benchmarks, so forcing should be a deliberate latency trade, not a default. The second is max_new_tokens: THINK traces are long relative to a 16,384-token budget, and truncating inside a reasoning block yields no answer at all rather than a short one.
SamplingShaping the output distribution
The repository ships two contradictory answers and you should know which you are following. generation_config.json sets do_sample true with temperature 0.8 and top_p 0.95, so any framework that reads that file will sample fairly loosely by default. Microsoft's own sample_inference.py in the same repository calls generate() with do_sample=False — greedy decoding, no temperature applied at all. The model card text gives no recommended sampling values, and the technical report was not read this session, so neither setting is stated as the endorsed one. The sensible reading: the shipped config is a general-purpose default, and greedy is what Microsoft used to demonstrate correctness. Start greedy for anything you are going to parse — coordinates, OCR output, extracted numbers — because this model's high-value outputs are the deterministic kind. Move to 0.8 with top_p 0.95 when you want the THINK trace to explore alternatives rather than commit to its first decomposition. There is no vendor guidance on whether to vary temperature and top_p together.
ReasoningThinking, effort and budgets
Yes, and it is unusually explicit. The model is a single system with two modes rather than two checkpoints. THINK mode emits a <think>...</think> block containing the reasoning, then the answer; NOTHINK mode tags a direct <nothink> response. Routing is normally automatic, driven by the system prompt printed on the model card, which instructs the model to pick a mode by complexity and confidence. You can force either by appending <think> or <nothink> to the generation template after the assistant turn marker. There is no numeric thinking budget and no effort ladder — you get a binary. Because you run the weights yourself, the thinking tokens are fully visible in your stream and cost you nothing but GPU time and context. Microsoft's own tables quantify the trade: forcing nothink drops ChartQA from 83.3 to 76.5 and MathVista-MINI from 75.2 to 68.7, while forcing think drops AI2D from 84.8 to 79.7 but lifts MathVerse-MINI from 44.9 to 53.1.
ToolsFunction calling and server tools
This is the field where the model diverges most from a general chat model. The card documents no function-calling schema, no tools array, no structured-output mode — there is nothing to enable and nothing to misconfigure. What it is built for instead is computer use: interpreting screen content, localising interactive GUI elements, and selecting actions. That capability is measured rather than declared, at 88.2 on ScreenSpot-V2, with ScreenSpot-Pro and V*Bench used internally during development. So the tool story here is inverted — the model is the perception layer inside your agent loop, and your own code holds the tool schema and does the clicking. If you need JSON out, enforce it with a grammar or constrained decoder in vLLM rather than expecting the model to honour a response_format. The failure mode people hit is dropping the system prompt: without it the THINK / NOTHINK routing is undefined, and a perception task can come back wrapped in a long unwanted reasoning trace.
CostPrice, caching, batching, what drives the bill
There is no per-token list price to quote, and that is the honest answer rather than a gap. The weights are MIT-licensed and the distribution channels Microsoft names are Hugging Face, GitHub and Microsoft Foundry. Azure's Phi pricing page renders its serverless table without figures for the models in question, and no serverless per-token rate for Phi-4-reasoning-vision-15B could be confirmed from a Microsoft source as of 2026-07-25 — Foundry deployment for this model appears to be managed-compute, billed as GPU hours rather than tokens. What actually drives your bill is therefore hardware. Microsoft reports testing on A6000, A100, H100 and B200, and recommends serving in bf16 on vLLM, which puts 15B bf16 weights at roughly 30GB before KV cache — a single 40GB or 80GB card. The 16,384-token window caps KV cache growth, which is a real cost advantage over long-context models. There is no caching discount and no batch discount because there is no vendor meter. The tokenizer has a 100,352-token vocabulary; images are the expensive input, consuming up to 3,600 visual tokens each against that 16,384 budget.
Where it runsSurfaces and availability
| Surface | Available | Notes |
|---|---|---|
| Microsoft Foundry | Yes | Named as a distribution channel on the model card, with a catalogue entry for Phi-4-Reasoning-Vision-15B |
| Hugging Face weights | Yes | microsoft/Phi-4-reasoning-vision-15B, MIT, ungated, seven safetensors shards |
| GitHub | Yes | Named on the model card as a public code channel |
| Self-hosting | Yes | Microsoft recommends vLLM >= 0.15.2 in bf16; transformers >= 4.57.1 and torch >= 2.7.1 with trust_remote_code |
| Hugging Face Inference | Unverified | No Microsoft statement about hosted inference for this repository was found |
| Amazon Bedrock | Unverified | Not named in any Microsoft source read this session |
| Google Vertex AI | Unverified | Not named in any Microsoft source read this session |
StrengthsWhat it is good at
- MIT licence on the weights — genuinely permissive, unlike the custom community licences attached to most open-weight releases at this size.
- 15B in bf16 fits a single 80GB accelerator, and Microsoft reports testing on A6000, A100, H100 and B200, so the deployment envelope is unusually broad for a vision reasoner.
- 88.2 on ScreenSpot-V2 makes it a credible perception layer for computer-use agents, a niche most small models score near zero on — Microsoft's own table puts gemma-3-12b-it at 3.5 on the same benchmark.
- Automatic THINK / NOTHINK routing beats both forced modes on most of Microsoft's published benchmarks, so the default path is also the strong path.
- Trained on 200 billion multimodal tokens over four days on 240 B200s — a small enough recipe that Microsoft published it, which makes the model unusually easy to reason about and to fine-tune against.
LimitsWhere it falls down
- There is no Phi-5. The site's row also attributes a 128K window to it — that figure matches
Phi-4-mini-instruct, whoseconfig.jsonsetsmax_position_embeddingsto 131072 with longrope scaling, and not this model. - 16,384 tokens is a small window in 2026, and a single high-resolution image can consume 3,600 of them before any text is read.
- English only. The card states plainly that the model is not intended to support multilingual use and that other languages may degrade.
- No tool-calling or structured-output contract is documented, so JSON must be enforced by your serving stack rather than requested from the model.
- Microsoft's own comparison tables show Qwen3-VL-32B-Instruct ahead of it on most benchmarks — MMMU 68.6 against 54.3, MathVision-MINI 54.3 against 36.2 — so the claim is efficiency at 15B, not frontier accuracy.
Against its neighboursHow it compares
Against Gemini 3 Flash the split is deployment, not capability: Flash is a hosted API you rent per token with a far larger window, while this is 30GB of MIT-licensed bf16 you put on your own card and run in an air-gapped network. If data residency or per-request cost at high volume is the binding constraint, that difference decides it. Against Claude Haiku 4.5, same shape of argument — a managed efficient-tier model with tool calling and structured output as first-class features, versus weights with neither but no meter attached. The comparison Microsoft itself invites is against Qwen3-VL-8B and Qwen3-VL-32B, both printed in its benchmark tables; the 8B is close on several rows and ahead on OCRBench and MMMU, and the 32B is ahead almost everywhere. Choose this one for GUI grounding, MIT terms, and a single-card footprint — not for topping a leaderboard.
Getting startedThe smallest call that works
vllm serve microsoft/Phi-4-reasoning-vision-15B --dtype bfloat16 --trust-remote-code
POST http://localhost:8000/v1/chat/completions
Content-Type: application/json
{
"model": "microsoft/Phi-4-reasoning-vision-15B",
"messages": [
{"role": "system", "content": "You are Phi, a multimodal model trained by Microsoft to help users."},
{"role": "user", "content": "Read the value shown on the gauge in this image."}
],
"temperature": 0.8,
"top_p": 0.95,
"max_tokens": 4096
}Microsoft states the vLLM floor and bf16 precision but prints no serve command; --trust-remote-code follows from config.json routing the model through modeling_phi4_visionr.py. Change the system prompt first — paste the full one from the model card so THINK / NOTHINK routing works. Then drop temperature to 0 for anything you parse.
SourcesWhere every claim above came from
- microsoft/Phi-4-reasoning-vision-15B model card — release date 4 March 2026, 15B, 16,384 token context, MIT, THINK / NOTHINK system prompt, training compute, benchmark tables
- Phi-4-reasoning-vision-15B config.json — hidden_size 5120, 40 layers, 40 heads, 10 KV heads, vocab 100352, rope_theta 500000, max_position_embeddings 32768, tokenizer_model_max_length 16384, SigLIP-2 vision tower
- Phi-4-reasoning-vision-15B generation_config.json — do_sample true, temperature 0.8, top_p 0.95, eos 100265, pad 100349, bos 100257
- Phi-4-reasoning-vision-15B sample_inference.py — Microsoft's own call uses trust_remote_code, bfloat16, max_new_tokens 4096 and do_sample False
- Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model — Microsoft Research, 200 billion multimodal training tokens, mid-fusion design, Foundry / Hugging Face / GitHub availability
- microsoft/phi-4 config.json — used to establish the shared backbone shape: hidden 5120, 40 layers, 40 heads, 10 KV heads, rope_theta 250000, max_position_embeddings 16384
- microsoft/Phi-4-mini-instruct config.json — max_position_embeddings 131072 with longrope scaling, the actual origin of the 128K figure in the site's row
- Phi Open Models — Microsoft Azure — product page listing the Phi family; no Phi-5, newest named are Phi-4-mini and Phi-4-multimodal
- Could not confirm: any model named Phi-5, from Hugging Face's microsoft organisation listing, Microsoft's Phi product page, or Azure documentation. The site's row for Phi-5 — including its 128K context figure — is not supported. Also unconfirmed: any serverless per-token price for Phi-4-reasoning-vision-15B on Azure or Microsoft Foundry, the model's knowledge cutoff as distinct from its training dates, a stated maximum output length, and availability on Bedrock, Vertex AI or Hugging Face Inference. The technical report PDF was not read this session.