Why this page existsA board that could not tell you when it had gone wrong
This site’s frontier board advertised a model that does not exist, called two API-only models open-weight, and printed a “longest context” superlative that was off by half. Every one of those was true-looking, aggregator-sourced, and sat there for days. Nothing on the page could tell a reader when a row had last been checked — so nothing could tell them it had gone wrong.
This page is the fix: a small register where every row carries the date it was checked and the method used, so it can be audited rather than believed. Aggregator error is not malice; it is what happens when a plausible name meets a plausible spec sheet and nobody spends the five minutes. The remedy is not to trust harder, it is to write down when you last looked.
A claim without a date and a method does not go on this page. Everything else here follows from that — including the fact that the register is short.
The nine entries are grouped by kind of wrongness rather than by vendor, because each kind fails differently and is caught by a different move. A model that does not exist and a model whose licence forbids shipping are not the same mistake, and looking for them the same way finds neither.
Section 02 · 2 entriesIt does not exist
The hardest class to catch, because a plausible name plus a plausible spec sheet reads exactly like a real release. Both of these were on this site’s own frontier board, and both sat there for days.
“Llama 5 — 600B open weights, 5M-token context, April 2026”
Repeated by aggregators and press-release syndication · checked 2026-07-25
What is true. There is no Llama 5. Not on Meta’s blog, not on developer.meta.com, and not under meta-llama on Hugging Face — whose newest repository of any kind was created 2025-04-28. Meta’s newest open weights are Llama 4 Scout and Maverick, April 2025, under the gated Llama 4 Community License. Its current frontier model is Muse Spark 1.1, announced 2026-07-09, and its weights are not released.
How it was checked. The meta-llama organisation on the Hugging Face models API, sorted by creation date — newest repo 2025-04-28, which settles it on its own. Then the ai.meta.com/blog post index for a Llama 5 entry: none.
“Phi-5 — small, open-weight, 128K context”
The 128K figure traces to a different model entirely · checked 2026-07-25
What is true. There is no Phi-5 — no such repository under microsoft on Hugging Face and no mention on Microsoft’s own Phi product page. The newest Phi is Phi-4-reasoning-vision-15B: 15B dense, MIT, released March 2026, with a 16,384-token window. That is a quarter of what the claim said, which changes whether it fits your job at all.
How it was checked. The microsoft HF organisation and the Azure Phi product page for existence. Then config.json on both candidates to find where the 128K came from: it is Phi-4-mini-instruct’s max_position_embeddings of 131072 with longrope scaling, borrowed onto a model that does not exist.
The 128K is the tell, and it generalises: a number that is correct for some model in the family is the signature of a fabricated row. Nobody invents 131072. It gets carried across from a sibling, and it is the thread worth pulling whenever a spec looks oddly specific for a model you cannot find.
Section 03 · 2 entriesThe weights are not there
The model is real. The download is not. This is the class that wastes the most time, because you only discover it after you have planned the deployment around owning the weights.
“Qwen3 Max — the leading open-weight family”
This site’s own frontier board said exactly this · checked 2026-07-26
What is true. No Qwen3-Max repository exists under the official Qwen organisation. It is commercial API only, priced in three tiers by input length. Licence: not applicable, because no weights were released. Alibaba ships plenty of open weights — this model is not among them, and the family’s reputation was doing the work the evidence should have done.
How it was checked. The official Qwen organisation on Hugging Face, searched for a Max repository: none. Then the vendor pricing page, which lists API tiers and no download.
“Kimi K3 — API and open-weight”
Announced as open. Not yet published. · checked 2026-07-26
What is true. No Kimi-K3 repository on the moonshotai organisation, and no licence published on any Moonshot page. An announcement is not a release, and until the weights are posted there is no licence to evaluate — so there is nothing to plan a self-host against.
How it was checked. The moonshotai HF organisation for a K3 repository: none. Then Moonshot’s own pages for licence text: none anywhere.
“Open” on a comparison board is read as I can take this in-house. That is an architecture decision, a budget line and sometimes a compliance answer. Getting it wrong does not send someone to a slightly worse model — it sends them down a path that has no end, and they find out at the point of download.
Section 04 · 3 entriesThe licence is not what “open” implies
Weights you really can download, under terms that decide whether you may ship. Compressed to a table, because the claim and the truth are one line each and the licence name is the whole story.
| The claim | The licence, actually | What it costs you | Checked |
|---|---|---|---|
| Falcon 3 — “fully open, permissive licence” | TII Falcon-LLM 2.0 | Not Apache, not OSI-approved. Ungated to download, but on TII’s terms — read them. “Permissive” was doing unearned work: ungated is not the same as free to ship. | 2026-07-25 |
| Command A — “open weights” | CC-BY-NC-4.0 | Non-commercial only, plus Cohere’s Acceptable Use Policy. You may evaluate; you may not ship. The weights genuinely are published, which is what makes this one easy to get wrong. | 2026-07-25 |
| Llama 4 — “open source” | Llama 4 Community | A custom commercial licence, and gated behind an access request. Neither half of “open source” survives contact with it. | 2026-07-25 |
The move that settles all three is the same, and it takes about fifteen seconds: read the licence field on the repository, not the word “open” in the announcement. The announcement is marketing copy about a release; the licence field is the thing your legal review will actually read.
Worth saying plainly, because it cuts the other way too: a marker that is wrong generously is still wrong. When this site audited the whole column rather than the row it had been sent to, two models were marked open that are not — and two were understated that are genuinely Apache 2.0 and MIT. Under-claiming hides a model somebody could have used.
Section 05 · 1 entryThe number is wrong
“Gemini 3.1 Pro — 2M context, the longest on the board”
Both this site’s fact file and its catalogue said 2M · checked 2026-07-26
What is true. 1M, not 2M. And the number is not the interesting part: crossing 200K input tokens reprices the entire request, not merely the tokens above the threshold. A budget built on the headline figure is wrong twice — once on how much fits, and once on what it costs when it does.
How it was checked. Google’s own model documentation for the context figure, and the vendor pricing page for the 200K threshold behaviour. The row now carries verified_by: vendor docs 2026-07-26.
“Longest context” is not a spec, it is a recommendation. A wrong spec misinforms; a wrong superlative decides — it is the cell someone sorts by, and the reason they pick a model without reading the rest of the row. It was off by half.
Section 06 · 1 entryWhere it does not run
Not a model claim — a capability claim, and the same failure shape. It is here to establish that this register is about checkable claims, not about models specifically.
“Agent Skills work anywhere Claude is served”
A reasonable assumption. Wrong on two major clouds. · checked 2026-07-25
What is true. Skills are absent on Amazon Bedrock and Google Vertex AI. They execute through the code-execution tool, and code execution is itself unsupported on both — so the entire path is missing rather than merely disabled. There is no flag to turn on, which is why waiting for one is the wrong plan. Note also that Amazon Bedrock and Claude Platform on AWS are two different products on the same cloud.
How it was checked. Already sourced on this site’s Agent Skills grid, with an asof date on the row itself — so this entry reads that row rather than restating it, and the two cannot drift apart.
The credibilityWe published five of these ourselves
A register that only corrects other people is marketing. So this one opens its own board: until 2026-07-30, this site’s frontier table advertised a model that does not exist, called two API-only models open-weight, and printed a “longest context” superlative that was off by half.
| Row | What this site said | What it says now |
|---|---|---|
| Llama 5 | Open-weight · 5M context · 600B dense | Muse Spark 1.1 · Closed API · weights not released |
| Phi-5 | Open-weight · 128K | Phi-4-reasoning-vision-15B · MIT · 16K |
| Qwen3 Max | Open-weight · “leading open-weight family” | Closed API · no weights released |
| Kimi K3 | Open-weight | API only · weights announced, not published |
| Gemini 3.1 Pro | “Longest context (2M)” | 1M · and the whole request reprices past 200K |
Each was true-looking, aggregator-sourced, and sat there for days. Eleven corrections shipped together on 2026-07-30, and the count matters more than any single row: the handoff naming this problem named one bad row. Checking the whole column found eleven, in both directions.
The Llama 5 row was stamped not vendor-confirmed on the day it was written. The honesty was already in the data. Nothing ever read it. Provenance you record but never query is decoration — and the same audit found that the one row with verified_by: None was, again, the one row with a real problem.
Publishing our own errors is not self-flagellation; it is the only thing that makes the other four entries worth believing. A page that claims to be right about other people, and says nothing about itself, has given you no way to check it — which is the exact failure this register exists to name.
The Agent Skills grid already solved this: every one of its pages ends with a freshness note read straight off the grid row, so the page and the table cannot disagree about when a claim was checked. The models board had no such binding, which is exactly why a false row survived eight days. Section 06 above is the proof it works — that entry reads its date from the grid rather than restating it.
What transfersCheck it yourself, in four moves
This is the method that found every entry above. It is four moves and it takes about five minutes per claim — which is the actual finding here, because five minutes is cheap and almost nobody spends it.
| Move | What it settles | What it looks like |
|---|---|---|
| 1 · Go to the org, not the model | Whether it exists at all | The vendor’s organisation on the Hugging Face models API, sorted by creation date. A newest-repo date older than the claimed release settles it without reading a word. |
| 2 · Read the licence field, not the word “open” | Whether you may ship it | The repository’s licence metadata and the organisation card — not the launch blog post. |
| 3 · Take numbers from config, not from prose | Context, parameters, precision | config.json — max_position_embeddings, any rope scaling, and the tokenizer’s own cap, which can be lower. |
| 4 · Write down what you could not confirm | Where the next reader should be careful | An explicit “could not confirm” list, in the artifact itself — not silence. |
Every model page on this site ends with an enumerated list of what could not be verified — max output, knowledge cutoff, whether reasoning tokens bill as output. Most AI writing states what it found and goes quiet about the rest, and a reader cannot tell the difference between checked and fine and never looked. Naming the gap is the cheapest honesty available and almost nobody does it.
Moves 1 and 3 are also the two that a machine can run unattended, which is where this goes next: a claim whose as_of stamp is older than its vendor’s newest release is a claim worth re-checking, and that comparison needs no judgement at all. Every row on this register reads its date from the same fact file the board reads, so the two cannot disagree about when a claim was last true.
Last full review of this register: 2026-07-30. The nine entries above were checked between 2026-07-25 and 2026-07-26; all five site-owned errors were corrected and deployed on 2026-07-30.
What changedWhat changed here
Updated this page A cheaper Opus 5.5 reportedly changes agent-facing behavior, so swapping the model string can pass smoke tests and still fail in production.
Add Claude Opus 5.5 to the Opus line and note that the cheaper model changed agent-facing behavior in ways smoke tests miss.
Three kinds of claim, strongest first. Signal runs every morning.