Answer customer questions from your product manuals
Customers ask questions your manuals already answer, but finding the right page takes time. This app searches the manuals, finds the passage that answers the question, and drafts a reply with that passage shown beside it.
PresenterOpens the private repo. Visible to admins only.
For the support deskCross-domain
Why it matters
Today's manual process, and the same job with the app
A support desk at a software company, working from its own product manuals.
✕Today's manual process
1Read the question and guess which manual might have the answer.
2Search the manuals page by page, hoping the right words were used.
3Write the reply from memory, hoping the quote is right.
4Miss the real answer and the customer gets an incomplete or wrong reply.
Every question searched manually
✓With the app
1The question is read and the manuals are searched for the passage that answers it.
2The best passages are found ranked, so the closest match comes first.
3The reply is drafted with the exact passage quoted beside it.
4If the answer isn't there it says so, instead of guessing.
Every answer arrives with its source
See it work
One real case, read by the app, step by step
A reader asks why SQLite can't be compared to client/server databases, and gets an answer with the passage it used shown beside it.
Answer customer questions from your product manualsReference appBuilt to be shaped to your process
4
1The question typed in by the reader.
2The top passage ranked highest among the passages the app searched.
3The answer drafted with its two source passages marked.
4The second source the other passage the answer's citation points to.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Answer customer questions from your product manuals
A small, forkable project that does one job end to end. Run twice for real over the same set, and every figure on these pages captured from those runs.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
Someone has a body of written material — product documentation, policies, contracts, runbooks — and a stream of questions whose answers are already in there somewhere. Today those questions are answered by a person who knows where to look, or by nobody. Reading forty documents to find the paragraph that answers one question — the kit retrieves the passages and drafts the answer with them quoted beside it, so the check is reading five passages instead of forty files.
Audience
A product manager or architect deciding whether retrieval-augmented generation is worth building for their own documents — and what it will cost, how it fails, and how they would know it works. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual documents
The corpus is 40 documents, 1.45 MB (docx 3 · html 12 · md 12 · pdf 5 · txt 8). Five formats deliberately, so extraction is exercised rather than assumed. It is SQLite's own public documentation, chosen because it is genuinely public domain and can be redistributed inside a public repository.
The corpus
The 40 documentsunder its source's terms — https://www.sqlite.org/ — fetched 2026-07-31, mapped URL by URL in data/corpus/SOURCES.md.
Where each came fromrecorded in the kit's own SOURCES.md, beside the corpus it describes.
Swap this folder for your own material and the kit is pointed at your documents. That is the whole change — there is no database to migrate.
The outcomeWhat a good result looks like
A reader asks a question in their own words. The system finds the five passages most likely to contain the answer, shows them, and drafts an answer that cites which passage it used. The evidence and the answer arrive together, so the check is reading five passages rather than forty files.
And when it cannot
It says so. On three of fifty test questions the retrieval step handed the model the wrong passages, and the model declined rather than inventing something. That is the behaviour you want, and it is still a failure of the pipeline — both are reported here rather than one of them.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
A release gate — a wrong answer must not ship — Exact substring It almost never waves a failure through: TNR 1.000 and 0.833. Over 100 adjudicated answers 14 of its 15 mistakes are false FAILURES, which waste review time rather than shipping a regression — but 1 wrong answer did get past it, so it is a strong gate and not an infallible one.
Cheap triage over high volume, where a false alarm is expensive — Token F1 TPR 1.000 on both models and free, so it almost never flags a good answer. Use it to shrink what a human or a judge has to read.
The criterion needs interpretation — tone, faithfulness, whether a paraphrase counts — LLM judge, zero-shot The only grader here that reads meaning. It is what caught the refusal the substring check scored correct.
You need to know whether it is hallucinating — Faithfulness It asks a different question from correctness — does the answer follow from the passages — so it catches invention that a correctness grader scores as right.
The data cannot leave your network — Any code grader All four code graders run in-process, cost nothing and send nothing. Three of them beat the null baseline.
Ordered or procedural answers, where sequence carries meaning — ROUGE-L It is the only code grader that rewards ORDER, via longest common subsequence.
The judge disagrees with you and you cannot tell who is right — Human review It is the only grader that is actually right, and the other six are calibrated against it. Adjudicating only the DISAGREEMENTS is the cheap way in — 15 rows here rather than 100 — but it cannot give either grader a rate, because the sample is chosen on the very thing being measured. On those 15 the substring check scores 0%; over all 100 it scores 85.0%.
At a glanceHow the whole thing runs
88–90%judged correct (LLM-as-judge) · 2 runs, no ordering
2,040 msp50, end to end
$0.21per 1,000 questions · Google Gemini 2.5 Flash-Lite
Run twice over the same set, for real, the last on 2026-08-01. Every figure on these pages was captured from those runs — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Answer customer questions from your product manuals14 steps · 4 questions · run once, for real · 2026-08-01
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Replace the files in data/corpus/ and run python -m src.index. Corpus lens →
When is this the wrong choice?
Avoid: Token F1 — TNR 0.333 on the deliberating tier means two of every three real failures get through. That is the case against the best-fitting scenario (“A release gate — a wrong answer must not ship”). 7 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Documents whose answer lives under a heading the splitter separates from it — FOUR of the five recorded retrieval failures are this (L26, L28, L31, L46: the right document was retrieved and the wrong chunk of it ranked), and the table of contents outranks the content because it is dense with the question's own words. The fifth, L32, is NOT this class — its document never surfaced at all. 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
not_extracted can never fire. Labels are validated against the extracted text and the taxonomy inspects that same text, so anything lost in extraction fails the label gate before it can be scored. 7 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 2 models on the fast tier and the deliberating tier, one provider, one key. Prompt lens →
A model graded this — who checked the grader?
A model does the grading here, so the score you are about to trust was itself re-judged — the adjudication run is published as its own page rather than summarised. the adjudication run →
And if it fits — what do I stand up?
4 artifacts with a stated home and a stated egress, and 4 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 1 question: clone (a fresh clone of this kit runs with nothing fetched).
Last verified 2026-08-01 — two models, one provider, one key — re-judged 2026-08-01. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — Verified from a cold clone on the day of the run: 40 documents, 668 chunks, and the committed retrieval numbers reproduced exactly, with no key configured.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
94.0%rows answered
2,040 msp50, end to end
3,753 msp95
0.1 minclone to first result
What the clock covers. end-to-end per answered row: retrieval plus the model call, nearest-rank over the 50 answered rows, the same convention evals/run.py uses for its model-only figures
Current processWhat it replaces
Reading forty documents to find the paragraph that answers one question — the kit retrieves the passages and drafts the answer with them quoted beside it, so the check is reading five passages instead of forty files.
Where it is not good enough
Three of fifty questions get no answer at all: the model correctly declines when retrieval hands it the wrong passages (L26, L31, L46), so coverage is 94%, not 100%. All five retrieval failures are the same shape — the right document comes back and the chunk that wins is its table of contents — which is a splitter problem this kit has not solved. Latency is a second to four seconds, fine for a person and too slow for anything interactive. And the 50-row set is too small to separate two models: judged 90% against 88% is one row.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
Run once, for real. Every figure on the page comes from this run.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
PROVIDER, BASE_URL, MODEL in .env. Two adapters ship: openai-compatible, which covers every provider speaking the OpenAI chat-completions shape plus any local server (Ollama, LM Studio, vLLM), and anthropic, which is its own shape. Adding one is a function and a dict entry.
extraction
src/extract.py
add a format. The corpus is deliberately five formats so this seam is exercised rather than asserted.
chunking
src/chunk.py
MAX is a hard ceiling, read by the UI rather than retyped. All five recorded failures live behind this seam.
retrieval
src/retrieve.py
keyword and embedding sit behind one call; the index records which built it, and dispatch is on what the index HAS, not on what .env asks for.
prompt assembly
src/prompt.py
change the wording and the parts change with it — the decomposition is built as the string is built.
Components
Component
File
Role
extract
src/extract.py
PDF/DOCX/HTML/MD/TXT to text; pypdf is the one dependency in the kit
boilerplate
src/boilerplate.py
drops furniture repeated ACROSS the corpus — 909 lines over 36 distinct patterns
chunk
src/chunk.py
splits to a hard 1800-char ceiling; 668 chunks, p50 865
index
src/index.py
writes chunks.json; deterministic, so a rebuild that changes the index is a real diff
retrieve
src/retrieve.py
stdlib TF-IDF, top-5, no key and no network
prompt
src/prompt.py
assembles the prompt AND its decomposition in one pass, so they cannot disagree
adapters
src/adapters/__init__.py
one completion call over raw HTTP; every provider returns the same shape including usage
app
src/app.py
stdlib http.server UI; holds no pipeline logic
run
evals/run.py
the eval harness: retrieval pass, answer pass, cause taxonomy
judge
evals/judge.py
LLM-as-judge for correctness, with a self-test that must go red first
Where it breaks at scale
Retrieval scores every chunk on every query in Python: 668 chunks is 17ms, and it is linear, so a corpus a hundred times larger is a second or two per query before the model is called. chunks.json is loaded whole into memory at startup. There is no incremental index — adding one document rebuilds everything (0.36s here, minutes at a hundred thousand documents). The honest ceiling is a corpus you can hold in RAM and rebuild while you wait; past that the answer is a real vector store, and the retrieval seam is where it goes.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
A real question from the labelled set, answered live. The five passages that won are shown beneath the answer with their scores, so the evidence and the answer are on one screen.successOpen full size →The recorded runs, shipped in the repo. This tab renders with no API key, which is what lets a forker see the thing work before spending anything.successOpen full size →Extraction and chunking — the two seams a query never touches. 40 documents across five formats, 668 chunks, and where the splitter cut each one.successOpen full size →First screen after python -m src.app. Corpus size, retriever and model are stated up front; nothing asks the reader for a key.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
L26, a recorded failure. The right document WAS retrieved — three of the five passages are from foreignkeys and marked EXPECTED — but the chunks that won hold its table of contents, not the instructions. The model declined rather than guessing, which is the correct behaviour and still a failure of the pipeline.failureOpen full size →
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
40documents
1.45 MiBdocx 3 · html 12 · md 12 · pdf 5 · txt 8
668chunks · p50 865 chars
$0.00index build · 0.36s
How it is cutWhat one chunk is
paragraph-greedy to a hard 1800-character ceiling, after corpus-wide boilerplate removal
The indexWhat the index build measured
There IS an index here, and this is the one kit of the fifteen that builds one: src/index.py writes 668 chunks to chunks.json in 0.36s, and it is a file on your disk rather than a service. Retrieval is stdlib TF-IDF over that file, top-5, with no key and no network — which is why the build costs $0.00. It is deterministic, so a rebuild that changes the index is a real diff rather than noise.
LicenceLicence
Public domain. SQLite's dedication names documentation as well as code, verified against sqlite.org/copyright.html on 2026-07-31. No attribution clause, no share-alike, no notice file.
Bring your ownBring your own documents
Replace the files in data/corpus/ and run python -m src.index. That is the whole procedure: the index is a file, it is deterministic, and the rebuild takes under a second for forty documents. The labelled set in data/labelled.jsonl is about THESE documents, so your accuracy numbers start meaningless until you write your own rows — evals/check_labels.py enforces six rules on them and will tell you what is wrong before you spend anything on a run.
What breaks it
Documents whose answer lives under a heading the splitter separates from it — FOUR of the five recorded retrieval failures are this (L26, L28, L31, L46: the right document was retrieved and the wrong chunk of it ranked), and the table of contents outranks the content because it is dense with the question's own words. The fifth, L32, is NOT this class — its document never surfaced at all.
Scanned or image-only PDFs: pypdf extracts no text and there is no OCR step.
A corpus where every document repeats the same nav furniture — handled here by dropping 909 repeated lines across the corpus, but that pass needs more than one document to work at all.
Any single passage longer than the 1800-character ceiling is cut, so an answer spanning the cut is not retrievable whole.
A file that DESCRIBES the corpus sitting inside it — SOURCES.md was silently indexed as a 41st document until 2026-08-01; src/index.py now names it.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — they were measured against the provider’s own counter, not apportioned by character share.
Token budgetWhere the tokens actually go
The five retrieved passages are 1,761 of 1,911 input tokens; the reader’s own question is 21. That ratio is the cost lesson, and it is why a per-query bill is driven by how much context you retrieve.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system prompt (incl. chat framing)
183
121
instruction
48
8
passage 1
1,445
307
passage 2
1,779
372
passage 3
1,816
354
passage 4
1,706
360
passage 5
1,749
368
question
110
21
Total
1,911
This is the cost lesson as arithmetic: of the 1,911 tokens assembled, 1,761 are passages — 92% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You answer questions using only the provided document passages.
If the passages do not contain the answer, say so plainly and do not guess.
Cite the passage number you used, like [2].
Answer the question using only these passages.
[1] whentouse (pdf)
Appropriate Uses For SQLite
Table Of Contents
SQLite is not directly comparable to client/server SQL database engines such as MySQL, Oracle, PostgreSQL,
or SQL Server since SQLite is trying to solve a different problem. Client/server SQL database engines strive to implement a shared repository of enterprise data. They emphasize
scalability, concurrency, centralization, and control. SQLite strives to provide local data storage for individual
applications and devices. SQLite emphasizes economy, efficiency, reliability, independence, and simplicity. SQLite does not compete with client/server databases. SQLite competes with fopen(). 1. Situations Where SQLite Works Well
Embedded devices and the internet of things
Because an SQLite database requires no administration, it works well in devices that must operate without
expert human support. SQLite is a good fit for use in cellphones, set-top boxes, televisions, game consoles,
cameras, watches, kitchen appliances, thermostats, automobiles, machine tools, airplanes, remote sensors,
drones, medical devices, and robots: the "internet of things". Client/server database engines are designed to live inside a lovingly-attended datacenter at the core of the
network. SQLite works there too, but SQLite also thrives at the edge of the network, fending for itself
while providing fast and reliable data services to applications that would otherwise have dodgy
connectivity.
[2] quirks (docx)
o has stumbled over some quirk of SQLite that is not mentioned here, please let the developers know by posting a brief message on the SQLite Forum. 2.
SQLite Is Embedded, Not Client-Server
Whenever comparing SQLite to other SQL database engines like SQL Server, PostgreSQL, MySQL, or Oracle, it is important first of all to realize that SQLite is not intended as a replacement or competitor to any of those systems. SQLite is serverless. There is no separate server process that manages the database. An application interacts with the database engine using function calls, not by sending messages to a separate process or thread. The fact that SQLite is embedded and serverless instead of being client/server is a feature, not a bug. Client/server databases like MySQL, PostgreSQL, SQL Server, Oracle, and others are an important component of modern systems. These systems solve an important problem. But SQLite solves a different problem. Both SQLite and client/server databases have their role. Developers who are comparing SQLite against other SQL database engines need to clearly understand this distinction. See the Appropriate Uses For SQLite document for additional information. 3. Flexible Typing
SQLite is flexible with regard to datatypes. Datatypes are advisory rather than mandatory. Some commentators say that SQLite is "weakly typed" and that other SQL databases are "strongly typed". We consider these terms to be inaccurate and even pejorative. We prefer to say that SQLite is "flexibly typed" and that other SQL database engines are "rigidly typed". See the Datatypes in SQLite document for a detailed discussion of the type system in SQLite. The key point is that SQLite is very forgiving of the type of data that you put into the database.
[3] faq (docx)
n many NFS implementations. You should avoid putting SQLite database files on NFS if multiple processes might try to access the file at the same time.
On Windows, Microsoft's documentation says that locking may not work under FAT filesystems if you are not running the Share.exe daemon. People who have a lot of experience with Windows tell me that file locking of network files is very buggy and is not dependable. If what they say is true, sharing an SQLite database between two or more Windows machines might cause unexpected problems. We are aware of no other embedded SQL database engine that supports as much concurrency as SQLite. SQLite allows multiple processes to have the database file open at once, and for multiple processes to read the database at once. When any process wants to write, it must lock the entire database file for the duration of its update. But that normally only takes a few milliseconds. Other processes just wait on the writer to finish then continue about their business. Other embedded SQL database engines typically only allow a single process to connect to the database at once. However, client/server database engines (such as PostgreSQL, MySQL, or Oracle) usually support a higher level of concurrency and allow multiple processes to be writing to the same database at the same time. This is possible in a client/server database because there is always a single well-controlled server process available to coordinate access. If your application has a need for a lot of concurrency, then you should consider using a client/server database. But experience suggests that most applications need much less concurrency than their designers imagine. When SQLite tries to access a file that is locked by another process, the default behavior is to return SQLITE_BUSY.
[4] whentouse (pdf)
exibility since
new columns and indices can be added without having to recode every query. Stand-in for an enterprise database during demos or testing
Client applications typically use a generic database interface that allows connections to various SQL
database engines. It makes good sense to include SQLite in the mix of supported databases and to
statically link the SQLite engine in with the client. That way the client program can be used standalone
with an SQLite data file for testing or for demonstrations. Education and Training
Because it is simple to setup and use (installation is trivial: just copy the sqlite3 or sqlite3.exe executable
to the target machine and run it) SQLite makes a good database engine for use in teaching SQL. Students
can easily create as many databases as they like and can email databases to the instructor for comments or
grading. For more advanced students who are interested in studying how an RDBMS is implemented, the
modular and well-commented and documented SQLite code can serve as a good basis. Experimental SQL language extensions
The simple, modular design of SQLite makes it a good platform for prototyping new, experimental
database language features or ideas. 2. Situations Where A Client/Server RDBMS May
Work Better
Client/Server Applications
If there are many client programs sending SQL to the same database over a network, then use a
client/server database engine instead of SQLite. SQLite will work over a network filesystem, but because
of the latency associated with most network filesystems, performance will not be great. Also, file locking
logic is buggy in many network filesystem implementations (on both Unix and Windows).
[5] whentouse (pdf)
ms) to analyze large datasets. Raw data can be imported from CSV files, then that data can be
sliced and diced to generate a myriad of summary reports.
More complex analysis can be done using
simple scripts written in Tcl or Python (both of which come with SQLite built-in) or in R or other
languages using readily available adaptors. Possible uses include website log analysis, sports statistics
analysis, compilation of programming metrics, and analysis of experimental results. Many bioinformatics
researchers use SQLite in this way. The same thing can be done with an enterprise client/server database, of course. The advantage of SQLite
is that it is easier to install and use and the resulting database is a single file that can be written to a USB
memory stick or emailed to a colleague. Cache for enterprise data
Many applications use SQLite as a cache of relevant content from an enterprise RDBMS. This reduces
latency, since most queries now occur against the local cache and avoid a network round-trip. It also
reduces the load on the network and on the central database server. And in many cases, it means that the
client-side application can continue operating during network outages. Server-side database
Systems designers report success using SQLite as a data store on server applications running in the
datacenter, or in other words, using SQLite as the underlying storage engine for an application-specific
database server. With this pattern, the overall system is still client/server: clients send requests to the server and get back
replies over the network. But instead of sending generic SQL and getting back raw table content, the client
requests and server responses are high-level and application-specific.
Question: Why is SQLite not directly comparable to client/server engines such as MySQL, Oracle or PostgreSQL?
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
SQLite is not directly comparable to client/server engines because it is trying to solve a different problem: SQLite provides local data storage for individual applications and devices, while client/server databases emphasize scalability, concurrency, centralization, and control [1]. Also, SQLite is embedded and serverless, not a client/server system [2].
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Answer customer questions from your product manuals — 50 questions drawn from 40 real documents. Two tiers of one model family answered, and every answer was then graded Seven different ways — because the grade is a measurement too, and most of what went wrong on this kit went wrong in the ruler rather than the model.
A further 91 questions were later authored and run as a probe — an attempt to build a harder set that could separate the graders better. They are not part of any score on this page, and what they found is under The specification.
The rulerHow it was graded
A MODEL DECIDES, and this is the only kit here where that is true — every other one grades with code. The judge's verdict is an opinion, so unlike an exact match it can be wrong in both directions, and the rates on this page move when it is. That is why seven rulers were run over the same 100 answers rather than one: the choice of ruler is itself a decision, and putting a grader that passes everything on the same table is what makes the others readable. The judge is the reference standard the other six are scored against, so nothing on this page can check the judge — only the adjudication can, and it was another model.
50questions
40source documents
2model tiers
100graded answers
7grading methods
91probe questions
MeasurementsWhat was measured
COUNTED49 / 50The right document reached the modelNo model in the loop. Reproduced from a cold clone with no API key.
COUNTED45 / 50The answer text was inside the promptThe model could have seen it. Not that it used it.
JUDGED45 · 44 / 50A second model accepted the answerA grader that accepts everything scores the same.
COUNTED37 · 39 / 50Answer contained the exact reference wordingA poor measure of correctness — it marks “Yes.” wrong. Kept as a grounding signal, not a score.
JUDGED85 / 100The substring check's verdict was rightMeasured against an adjudication of all 100 answers, including the 85 the graders already agreed on — TPR 84.3%, TNR 90.9%. The adjudicator was another model, not a person.
NOT YET KNOWN—A person confirmed the judge got it rightAll 100 answers have now been adjudicated and the judge matched every one — but the adjudicator was another model, so this row stays empty. The judge is the reference standard the other six graders are scored against, so nothing else here can fill it in.
The adjudication behind the substring figure — all 100 answers, the confusion matrix and the rows the graders fought over — is on Judge validation.
Three words carry this page: Counted is deterministic and nobody's opinion, Judged is a model's verdict that moves if the grader is wrong, and Not yet known is printed blank rather than filled with something plausible. The row most worth having is currently the empty one.
How the method was validated
Adjudicated. All 100 answers were given an independent verdict — not only the 15 the graders disputed — and the judge matched that adjudication on every row, 0 false passes and 0 false failures. THE ADJUDICATOR WAS ANOTHER MODEL, so this validates the judge against a second system built the same way and is not human ground truth; the method's own standard asks for a person and that has not happened. Its rates are therefore published as UNKNOWN rather than as 100%. See build/measured/judge-validation.json and the Judge validation page.
155output tokens · the fast tier · 2,023 ms p50
298output tokens · the deliberating tier · 4,642 ms p50
Two tiers, one API key
Not two vendors and not a vendor comparison — which is the whole reason the two-model rule could be met without signing up to a second provider. Both tiers send the same retrieved context; the slower one writes about twice as much back, waits 2.3× as long, and lands one row apart on 50. This set does not separate them on quality and no ranking should be read from it — which is not an opinion about the models but something the second run below demonstrated.
Run it twiceThe same set, run again
Two runs, and the ordering reversed. Same fifty questions, the same index, the same judge model and the same prompt: the fast tier finished ahead on the first run and behind on the second, by exactly one row each way. Neither ordering is published as a result, because neither is one. Two tiers a single row apart on fifty questions are not distinguishable on this set, and it took a second run to show it.
Run date
the fast tier
the deliberating tier
2026-08-01
90.0% r001-flash
88.0% r001-pro
2026-08-02
88.0% r003-flash
90.0% r003-pro
judged correct (LLM-as-judge) — The two runs are not averaged into one figure. An average of 90.0 and 88.0 is 89.0, which reads as a measurement nobody took and hides the only thing this pair actually established — that the gap is smaller than the noise. The runs are printed as runs.
What did not move
Everything deterministic reproduced to the digit across all five runs — document in top-5 0.980, answer present in the prompt 0.900, and the same five rows failing by name (L26, L28, L31, L32, L46). The instrument is steady. It is the fifty questions that cannot separate these two tiers, which is the same limit the null-grader row reports one section below and the reason the harder set was attempted at all.
Grading costWhat it costs
Every dollar published on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate. It is a projection onto a rate card, not a bill anyone paid. Token counts belong to this pipeline — top_k, chunk size, prompt — and hold wherever you run it; the dollars belong to whoever you buy from, which is why the card is named beside every figure.
Priced at
Per 1M in / out
One question
1,000 questions
Share that is the prompt
Google Gemini 2.5 Flash-Lite the cheapest model Google publishes a rate for
$0.10 / $0.40
$0.000211
$0.21
71%
Amazon Nova Pro on Bedrock a frontier tier, for the other end of the range
$0.80 / $3.20
$0.001690
$1.69
71%
Same work, 8× the bill
The same questions, the same tokens — only the rate card changed. And on either card about 71% of what you pay is the prompt this pipeline sends, not the answer it writes.
That is the one cost lever worth touching, and it is top_k — how many passages you retrieve per question.
Rates checked 2026-08-01. A cheaper American model than Gemini 2.5 Flash-Lite almost certainly exists — Amazon Nova Micro is widely reported at $0.035 and $0.14 per million. AWS's own pricing table would not render to an automated fetch on three attempts, so it is named here and left unpriced rather than quoted from a comparison site.
What grading adds
Grading IS priced here, on the same cards, from the judge's own recorded tokens — 337 in and 246 out per row on the fast tier, over 54 calls for 50 rows because an unreadable verdict escalates its ceiling. An earlier version of this page said the judge's tokens had never been recorded. They had: judge_input_tokens_total sits in every result file, and the claim was wrong for about an hour before a reader asked.
Grading unitWhat the grading figure prices
tokensgrading cost, as measured
GRADING COSTS MONEY HERE, AND ON EVERY OTHER KIT IT IS FREE. The unit is tokens because the grader is itself a model call: 337 in and 246 out per row on the fast tier, over 54 calls for 50 rows, since an unreadable verdict escalates to a higher ceiling. So the promise made elsewhere on this site — that a fresh clone can re-score the committed records for $0.00 — does NOT hold for the judge on this kit. What is free here is re-running the six code graders over the recorded answers; re-running the judge is a fresh bill.
The gradersSeven ways to grade
The judge passes about 90% of these rows, so a grader that returns “correct” to everything already agrees with it 90% of the time. That is the number every grader has to beat, and two of them do not.
The centre line is a grader that passes everything, not zero. Four beat it. Two tie it by construction. One scores below it — the exact-substring check that this kit originally graded correctness with.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Always “correct” the null grader — passes everything
$0.00
no
yes
the fast tier 90.0% · the deliberating tier 88.0%
Exact substring Does the answer contain a required fragment, verbatim?
$0.00
no
yes
the fast tier 84.0%(-6.0) · the deliberating tier 86.0%(-2.0)
Token F1 How much of the reference wording overlaps the answer, in both directions?
$0.00
no
yes
the fast tier 96.0%(+6.0) · the deliberating tier 92.0%(+4.0)
Reference recall How much of the reference answer's content made it into the answer?
$0.00
no
yes
the fast tier 96.0%(+6.0) · the deliberating tier 92.0%(+4.0)
ROUGE-L Does the answer follow the reference's wording in the same ORDER?
$0.00
no
yes
the fast tier 94.0%(+4.0) · the deliberating tier 90.0%(+2.0)
LLM judge, zero-shot Is the answer correct, judged on meaning rather than wording?
$0.36
yes — every row
no
the fast tier 90.0% · the deliberating tier 88.0%
LLM judge, few-shot The same question, with two worked examples in the prompt.
$0.36
yes — every row
no
the fast tier 90.0% · the deliberating tier 88.0%
Faithfulness Does the answer follow from the passages it was given — regardless of whether it is correct?
$0.36
yes — every row
no
the fast tier 100.0% · the deliberating tier 98.0%
Every row opens its own page: how the test was run, the input on one real row, the prompt or formula with every iteration of it, and the analysis.
Why this set cannot separate them
Every winning threshold landed at the BOTTOM of its sweep, which means these metrics agree mostly by passing almost everything, like the null grader. This labelled set cannot really separate them. The honest fix is a harder, balanced set — not a better metric.
Set limitationsWhat this set cannot show
Of 50 rows, 43 are judged correct on BOTH models. Those carry no signal about grader quality — every grader passes them, so they inflate agreement and separate nothing. Three rows are correct on exactly one model. Only FOUR (L26, L31, L41, L46) are wrong on both, and those four are the entire discriminating power of the set.
This is why every winning threshold landed at the bottom of its sweep, and why TNR moves by 0.17-0.20 when a single row flips. The set is not too small — it is too EASY, and adding more of the same rows would not help.
The specification
Target roughly half negatives. TNR on 25 negatives has an error bar a fifth the width of TNR on 5, and TNR is the metric an operator actually alarms on.
Sourcing the negatives from the FAILURE TAXONOMY is the right mechanism and the wrong rate. Retrieval failure does produce a wrong answer -- four of the five negatives found below are exactly that -- but retrieval puts the right document in the top five 96.7% of the time, so questions that defeat it have to be manufactured, and a set of manufactured retrieval failures measures the retriever rather than the graders.
Keep the existing 50 rows UNCHANGED as their own set. Changing labels invalidates every recorded run against them, and the published 90.0/88.0 are cited on four pages.
Add the new rows as a SECOND set with its own id prefix, and report both. A harder set is a different instrument, not a correction to this one.
This specification was attempted on 91 authored rows, and it does not produce a balanced set
91 questions were authored across 39 of the 40 documents, every answer fragment verbatim in the extracted text and no longer than four words, then answered by the fast tier and graded by the reference judge. Roughly half negatives was the target. 5 of 91 came back negative -- 5.5%. Four of those five are the rows where retrieval failed and the fifth is an empty answer. The pipeline is simply right about this corpus most of the time, so the negative class cannot be authored into existence; it would have to be manufactured, and that measures something else. Worse, the new questions are EASIER: a grader that passes everything agrees with the judge 94.5% of the time here against 90.0% on the original fifty.
The attempt above cost $0.025 for 91 rows on one tier plus judging, against the $0.06 estimated for 50 rows on two. The balanced set still does not exist, and this is now a measured gap rather than an untried plan: authoring harder questions is not the lever. A negative class big enough to alarm on needs a harder TASK -- a corpus with genuine ambiguity, or questions spanning two documents -- not a longer list of questions about this one.
Choosing oneWhich one to use
The comparison is only useful if it ends in a choice. The answer is conditional — and it is not the grader with the best agreement score.
It almost never waves a failure through: TNR 1.000 and 0.833. Over 100 adjudicated answers 14 of its 15 mistakes are false FAILURES, which waste review time rather than shipping a regression — but 1 wrong answer did get past it, so it is a strong gate and not an infallible one.
Token F1 — TNR 0.333 on the deliberating tier means two of every three real failures get through.
Cheap triage over high volume, where a false alarm is expensive
It asks a different question from correctness — does the answer follow from the passages — so it catches invention that a correctness grader scores as right.
Reading it as accuracy. An answer can be faithful to passages that were wrong.
It is the only code grader that rewards ORDER, via longest common subsequence.
Using it on one-clause factual answers, where order is noise and it is Token F1 with a penalty.
The judge disagrees with you and you cannot tell who is right
Human review
It is the only grader that is actually right, and the other six are calibrated against it. Adjudicating only the DISAGREEMENTS is the cheap way in — 15 rows here rather than 100 — but it cannot give either grader a rate, because the sample is chosen on the very thing being measured. On those 15 the substring check scores 0%; over all 100 it scores 85.0%.
Skipping it and quoting a judge accuracy as though it were truth.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
bad_ranking
the answer never reached the prompt
5
L26 — asked 'How do I turn on foreign key support in SQLite?'. Wanted foreignkeys; the passages that won were foreignkeys, foreignkeys, faq. The model answered: 'The provided passages do not give the steps to enable foreign key support. They only mention that…
answered_wrong
evidence was present, the answer was still wrong
8
L21 — asked 'How does WAL mode change the blocking relationship between readers and writers?'. Wanted wal; the passages that won were wal, wal, wal. The model answered: 'In WAL mode, readers and writers do not block each other; they can run at the same time…
dropped_in_chunking
extracted, but no chunk kept the answer whole
0
Zero rows. Not a claim that it cannot happen — it did not happen on these 50.
not_extracted
the text never survived the file
0
Zero rows, and the zero is UNINFORMATIVE — see could_not_verify. This category cannot fire by construction.
Technique choiceHow the grader was built
An LLM judge is one evaluation technique. The other five — deterministic, human review, weighted rubric, error analysis and hybrid — and the dashboard each one needs are on Evals, and a worked judge walkthrough is on the Evals dashboard.
Correctness here is judged by a model, and that is a decision with a history. The original grader marked an answer correct if it contained a required fragment. That is cheap, deterministic and needs no key — and it produced four different wrong numbers, each discovered only after fixing the last: table-of-contents fragments, then fragment length, then terseness, then paraphrase rigidity.
The tell was not the score moving. It was the ranking inverting. Noise moves a number; an artifact reorders the comparison the page exists to make. At that point the honest move is to change the instrument, not to relabel a fifth time.
Three things keep the replacement honest, and all three are checkable:
The judge never sees the required fragment.It gets the question, a reference answer and the candidate. Showing it the fragment would rebuild the instrument being replaced.
It was proved RED before it was trusted green.A self-test sends four fluent, confident wrong answers — including one that is true but answers a different question — plus two correct paraphrases that must pass. Six of six behaved as required. A grader that cannot fail anything reports 100% and means nothing.
The substring check still runsrenamed to what it actually measures: grounded. It fails in the opposite direction — it cannot be talked round by a confident wrong answer, and it cannot recognise a paraphrase. Every disagreement between the two is recorded.
One more thing worth stating, because it argues against the judge: a token ceiling covers thinking, not just output. A 400-token cap on a one-sentence verdict lost rows to empty and truncated replies, and re-rolling at the same ceiling reproduced the failure exactly. The ceiling now escalates, and only an unreadable verdict is retried — retrying a “false” would be shopping for the answer you want.
The full run log ships in the kit repo at results/judge-run-2026-08-01.log: the self-test block, every disagreement with the judge’s reason, and the row where the substring grader scored a refusal as correct because the model quoted the passage while declining.
What we could NOT verify
not_extracted can never fire. Labels are validated against the extracted text and the taxonomy inspects that same text, so anything lost in extraction fails the label gate before it can be scored. Its count of 0 says nothing about extraction quality. Closing it means re-authoring the labels against the SOURCE documents, which would invalidate every recorded run; for kit #1 it is stated rather than closed.
The two models are one row apart on fifty, and across two runs they swapped places — judged 90.0% / 88.0% on 2026-08-01 and 88.0% / 90.0% on 2026-08-02, on the same questions, index, judge and prompt. This dataset does NOT separate them on quality and no ranking should be read from it, in either direction. Every earlier version of the grader told a different story (86-vs-74, then 78-vs-74), which is why the instrument was replaced rather than tuned a fifth time.
judge_accuracy is a model's judgement, not ground truth. It never sees the answer_contains fragment, and it was proved red on four fluent wrong answers before use — but it can still be talked round by a confident wrong answer. The substring signal is kept as grounded because it fails in the opposite direction: it cannot be persuaded, and it cannot recognise a paraphrase.
The deliberating tier graded its OWN answers as well as the fast tier's. A model grading itself is a caveat, not a disqualification, and it is stated here rather than buried: on its own 50 rows it disagreed with the substring signal 7 times and overturned itself once (L27, where it marked a refusal wrong that the substring check had scored correct).
Latency was measured on one machine, on one network, against one provider's endpoint on one morning. It is not a property of the model.
Coverage counts a refusal as not-answered using a fixed list of refusal phrasings. A model that declines in wording outside that list would be counted as answering.
Semantic similarity (cosine over embeddings) is the one grader on the roster that was NOT run. The provider this kit's automation is allowed to use returns HTTP 404 for an embeddings endpoint, checked 2026-08-01, so there is no way to compute it here without introducing a second vendor. It is left unrun and named rather than substituted.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 2.5 Flash-Lite
Amazon Nova Pro on Bedrock
the fast tier
1,492
155
2,023 ms
$0.000211
$0.001690
the deliberating tier
1,413
298
4,642 ms
$0.000260
$0.002084
The token counts are measured on a real run and belong to this pipeline — top_k, chunk size, prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-01. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same question, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingGrading is a second bill on the same traffic
GRADING IS A SECOND BILL ON THE SAME TRAFFIC, AND IT IS ABOUT 39 PERCENT OF THE COMBINED ONE. The judge reads a question, a reference answer and a candidate — roughly 337 input tokens against the 1,492 the pipeline sends — so priced on a single card it costs LESS per row than answering does, not more. It still costs a second call, and its ceiling escalates on an unreadable verdict, so it made 54 and 56 calls for 50 rows rather than 50. A forker re-running the eval pays it ON TOP of the pipeline, and a code grader makes the column zero. Note what changed here: this page used to say grading cost MORE than answering, which was true only because the judge ran on the dearer of two tiers. On one rate card that reverses.
Cost driversWhat actually moves the bill
RETRIEVED CONTEXT, and it is not close. On the worked example the five passages are 1,761 of 1,911 input tokens; the user's question is 21. Doubling top_k very nearly doubles the bill, and the reader's question length barely moves it.
top_k — 5 here. It is one constant in src/retrieve.py and it is the largest cost lever in the kit.
The chunk ceiling (1800 chars). Bigger chunks mean fewer, larger passages; the product of k and chunk size is what you actually pay for.
Output length, which is the only cost difference between the two tiers once both are priced on the SAME rate card. The deliberating tier writes about twice as much — 298 tokens against 155 — for a judged score that is within one row of the fast tier's and that swapped places with it between two runs. Bigger gaps than that between two models are usually the rate card talking, not the workload.
Output length. The deliberating tier wrote ~298 tokens per answer to the fast tier's ~155, and output is priced at 2x input on both.
Grading, if you re-run the eval — see cost_of_evaluation_usd. It is the line people forget and it is larger than the pipeline itself.
Your volumeWhat it costs at your volume
Linear. Nothing here batches and nothing amortises: ten times the questions is ten times the bill, because each query retrieves different passages and so shares no cacheable prefix beyond the 121 tokens of system prompt and chat framing. The index is built once and costs nothing to serve, so corpus growth raises latency and memory before it raises the bill.
Where pricing changes shape
The provider used here prices a cache HIT at roughly 1/120th of a miss on the deliberating tier. This pipeline almost never hits: the retrieved passages sit at the FRONT of the prompt and differ per question, so the shared prefix is 121 tokens. Moving the passages AFTER a large fixed preamble is a real optimisation and the kit does not do it.
Both models are 1M context, so nothing here approaches a context-tier boundary — the worked prompt is 1,911 tokens, about 0.2% of the window. A kit whose prompt grew past a provider's repricing threshold would see its whole request reprice, not just the excess.
Switching the retriever from keyword to embedding adds a per-query embedding call AND a one-off index build cost that is $0.00 today.
Your return, with your numbers
VolumeNot assumed. Cost is linear, so a reader multiplies the per-question figure on their chosen rate card by their own query volume. On the cheapest American card that is about 21 cents per thousand questions.
What it replacesTime spent searching a document set by hand for the passage that answers a question, plus the reading to confirm it.
Time saved per itemNot measured. The kit measures its own latency (p50 2040ms end to end) and cannot measure how long the same question takes a person against your documents — that number is yours, and multiplying by your labour cost is what turns these figures into a return.
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
This kit runs on the cheapest capable model available, deliberately. The whole programme is self-funded community work, and every figure on this page was produced for about eleven cents — which is the point: the pipeline, the labelled set, the error taxonomy and the grading discipline are what transfer, and none of them depend on paying frontier prices to learn. What a run costs YOU depends on the model you already hold a key for, so the projection below prices the same measured workload across the models you are more likely to be running in production.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
179,346input tokens · this run
48,175output tokens
$0.093what it actually cost
The exact work behind every number on this page: 50 labelled questions answered by two models, and all 100 of those answers graded by the LLM-as-judge.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gemini-3-flash
Google
$0.106
$0.234
$1.21
2026-07-01
claude-haiku-4-5
Anthropic
$0.192
$0.420
$2.27
2026-07-22
gpt-5-6-luna
OpenAI
$0.212
$0.468
$2.42
2026-07-22
llama-5
Meta
$0.199
$0.429
$2.52
2026-07-25
grok-4-5
xAI
$0.303
$0.648
$3.91
2026-07-22
gemini-3-1-pro
Google
$0.423
$0.937
$4.84
2026-07-26
gpt-5-6-terra
OpenAI
$0.529
$1.171
$6.05
2026-07-22
claude-sonnet-5
Anthropic
$0.575
$1.261
$6.80
2026-07-22
claude-opus-5
Anthropic
$0.958
$2.101
$11.33
2026-07-25
claude-opus-4-8
Anthropic
$0.958
$2.101
$11.33
2026-07-25
gpt-5-6-sol
OpenAI
$1.059
$2.342
$12.11
2026-07-22
claude-fable-5
Anthropic
$1.917
$4.202
$22.67
2026-07-22
Read this against the numbers above
OUTPUT LENGTH IS THE WEAK HALF. 155 output tokens per query is what one model wrote; a terser or more verbose model moves that number and moves the estimate with it. Input dominates the bill here, which is what keeps the ranking meaningful even so.
Rates are read from build/facts/models.json, each row carrying its own as_of and source. A price is the fastest-ageing fact on this page — check the date beside it.
Cost is not quality. Nothing here says a dearer model would score better on this set: the two models actually measured finished within one row of each other on fifty questions, and which of them was ahead changed between the first run and the second.
No volume discount, committed-use rate, batch tier or cache hit is modelled. Each of those moves real enterprise pricing and none of them are on a public rate card.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
10 modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
src/extract.pyextract — a swap seam
PDF/DOCX/HTML/MD/TXT to text; pypdf is the one dependency in the kit
You change it to: add a format. The corpus is deliberately five formats so this seam is exercised rather than asserted.
src/extract.py
# SEAM 2 — text out of a document. Add a format by adding one entry to EXTRACTORS.
class HtmlText(HTMLParser):
def from_html(path, md=False):
def from_plain(path):
def from_docx(path):
def from_pdf(path):
EXTRACTORS = {
def extract(path):
src/boilerplate.pyboilerplate
drops furniture repeated ACROSS the corpus — 909 lines over 36 distinct patterns
src/boilerplate.py
# Drop the navigation, headers and footers that every document in a corpus repeats.
SHARE = 0.5
MAX_LEN = 160
MIN_DOCS = 3
SMALL = 5
def _repeated(texts):
def detect(docs, formats=None):
def strip(docs, formats=None):
src/chunk.pychunk — a swap seam
splits to a hard 1800-char ceiling; 668 chunks, p50 865
You change it to: MAX is a hard ceiling, read by the UI rather than retyped. All five recorded failures live behind this seam.
src/chunk.py
# SEAM 3 — how a document becomes retrievable pieces.
TARGET = 900
OVERLAP = 150
MIN = 120
MAX = 1800
def _cap(part, ceiling=None):
def split(text):
def chunk_document(doc_id, text, fmt):
src/index.pyindex
writes chunks.json; deterministic, so a rebuild that changes the index is a real diff
src/index.py
# Build and load the index. A file, not a server.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CORPUS = os.path.join(HERE, "data", "corpus")
INDEX = os.path.join(HERE, "data", "index")
CHUNKS = os.path.join(INDEX, "chunks.json")
VECTORS = os.path.join(INDEX, "vectors.json")
NOT_CORPUS = {"SOURCES.md", "README.md"}
def build(corpus_dir=CORPUS, verbose=True):
def load():
src/retrieve.pyretrieve — a swap seam
stdlib TF-IDF, top-5, no key and no network
You change it to: keyword and embedding sit behind one call; the index records which built it, and dispatch is on what the index HAS, not on what .env asks for.
src/retrieve.py
# SEAM 4 — which passages the model gets to see.
TOP_K = 5
def tokenize(s):
def build_keyword_index(chunks):
def keyword(question, chunks, idf, k=TOP_K):
def cosine(a, b):
def embedding(question, chunks, vectors, embed_fn, k=TOP_K):
def retrieve(question, index, cfg=None, k=TOP_K):
src/prompt.pyprompt — a swap seam
assembles the prompt AND its decomposition in one pass, so they cannot disagree
You change it to: change the wording and the parts change with it — the decomposition is built as the string is built.
src/prompt.py
# SEAM 5 — what the model actually receives.
SYSTEM = (
def assemble(question, passages):
src/adapters/__init__.pyadapters — a swap seam
one completion call over raw HTTP; every provider returns the same shape including usage
You change it to: PROVIDER, BASE_URL, MODEL in .env. Two adapters ship: openai-compatible, which covers every provider speaking the OpenAI chat-completions shape plus any local server (Ollama, LM Studio, vLLM), and anthropic, which is its own shape. Adding one is a function and a dict entry.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
def complete(cfg, system, user, max_tokens=1024):
src/app.pyapp
stdlib http.server UI; holds no pipeline logic
src/app.py
# The UI server. One command, no dependency, no build step, and it runs with no API key.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
RESULTS = os.path.join(HERE, "results")
LABELS = os.path.join(HERE, "data", "labelled.jsonl")
PORT = int(os.environ.get("PORT", "8765"))
STATE = {}
def _norm(s):
def load_state():
def api_status():
evals/run.pyrun
the eval harness: retrieval pass, answer pass, cause taxonomy
evals/run.py
# The eval harness. Scores the labelled set and classifies every failure.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
def _present(fragments, text):
def classify(row, hits, index, docs):
def retrieval_pass(rows, index, docs, k):
def answer_pass(rows, index, results, cfg, k, limit=None):
def summarise(results, k):
def main():
evals/judge.pyjudge
LLM-as-judge for correctness, with a self-test that must go red first
evals/judge.py
# The correctness grader. Reads a completed run and re-scores every answer with a model.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
VERDICT_TOKEN_SCALE = (400, 1200, 3000)
SYSTEM = (
def _read_verdict(text):
def _verdict(cfg, question, reference, candidate):
SELF_TEST = [
def self_test(cfg):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
src/extract.pyPDF/DOCX/HTML/MD/TXT to text; pypdf is the one dependency in the kit A swap seam.
src/boilerplate.pydrops furniture repeated ACROSS the corpus — 909 lines over 36 distinct patterns
src/chunk.pysplits to a hard 1800-char ceiling; 668 chunks, p50 865 A swap seam.
src/index.pywrites chunks.json; deterministic, so a rebuild that changes the index is a real diff
src/retrieve.pystdlib TF-IDF, top-5, no key and no network A swap seam.
src/prompt.pyassembles the prompt AND its decomposition in one pass, so they cannot disagree A swap seam.
src/adapters/__init__.pyone completion call over raw HTTP; every provider returns the same shape including usage A swap seam.
evals/run.pythe eval harness: retrieval pass, answer pass, cause taxonomy
evals/judge.pyLLM-as-judge for correctness, with a self-test that must go red first
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 1492 input and 155 output tokens per question at top-k 5, of which about 150 tokens are fixed (system prompt, chat framing, the question itself) and the rest is retrieved context. Change the retrieval depth and the input scales with it. Change the volume and the bill scales linearly, because nothing here batches.
The judge options add nothing to this total, and that is not because they are free
They are the opposite of free — on the run behind this kit, grading cost more than answering did. But the judge recorded only its dollar cost and never its token counts, so there is no measured quantity to multiply by the rate cards above, and converting the old figure would mean quoting a price list this page no longer names. Read a judged total here as the pipeline only. Instrumenting the judge is the fix and it is next.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Questions/day
Per day
Per month
Per year
Model and top-k follow the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Retrieval is free here and stops being free at scale.This kit scores every chunk in Python; a corpus a hundred times larger needs a vector store, which is infrastructure this arithmetic does not price.
Grading IS included now, and it is the surprise.An LLM judge costs more per row than answering does, so at full sampling it more than doubles the bill — and sends every row to a vendor. A code grader costs nothing and sends nothing. That is the control, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
A local, single-user tool over public documents. Binds 127.0.0.1, no auth, no rate limit, no session — correct for a demo on your own machine and wrong for anything else.
Read from the environment only, never written into the repo, never requested from a reader on any surface.
The experimentWe attacked it — two gates
An indirect prompt injection has to clear two gates, and they behave differently. Both were run for real on 2026-08-01, keyword retriever, top_k 5, 41-document scratch corpus.
Gate
Payload dressed as a doc page
Payload written to win
be retrieved
A payload dressed as an ordinary documentation page was NOT retrieved. The genuine lang_vacuum documentation took 4 of the 5 slots and nothing reached the model.
The same instruction in a document that repeats the query's own terms took RANK 1 of 5.
be obeyed
Nothing to obey — the payload never reached the prompt.
The injected instruction reached the assembled prompt VERBATIM. There is no filter between retrieval and the model. That is the finding.
The naive payload lost the retrieval race outright. Rewritten to repeat the query’s own terms — the oldest trick there is — the same instruction took rank 1.
The resultHijack failed, poisoning worked
0 of 10obeyed the injected instruction
5 of 10took CONTENT from the hostile document
5trials per model, two models
Neither model ever replied with the canary. But 5 of 10 answers repeated “defragment” — a claim that exists only in the hostile document and nowhere in SQLite's real documentation — and cited it as a passage.
Read this twice
The pipeline has no filter between retrieval and the model. Once a document is retrieved its text is in the prompt verbatim. A model declining to follow an instruction is not a defence — it is one vendor’s behaviour on one day, and it did not stop the answer being poisoned.
HonestyWhat this does not prove
One provider, one day, one payload, ten samples. A refusal rate is not a defence.
No test of a payload written to survive a semantic retriever rather than a keyword one.
Nothing was tested against a corpus containing private or personal data.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
You answer questions using only the provided document passages.
If the passages do not contain the answer, say so plainly and do not guess.
Cite the passage number you used, like [2].
src/prompt.py — it is the system prompt, and it is the only guardrail in the kit.
EvidenceDoes it hold?
What
Measured
Declined rather than inventing
On 3 of 50 questions retrieval handed the model the wrong passages and it said so instead of answering.
Stayed inside the evidence
Faithfulness judged 100.0% and 98.0% supported by the retrieved passages, red-proved on 4 seeded cases first.
A refusal the INSTRUMENT misread
L27: the model refused while quoting the passage, and the substring grader scored it CORRECT. The guardrail worked; the ruler did not.
The limitWhat a guardrail is not
A prompt rule is a request, not enforcement — nothing rejects an answer that ignores it.
No output filter, no schema validation on the answer, no PII redaction.
No defence against a hostile passage: the injection run put an instruction into the prompt verbatim and 5 of 10 answers took content from it.
No rate limit and no authentication in front of any of it.
WatchedWhat is watched, and why that one
5runs recorded
3 / 3deterministic metrics exact
+2.5%largest move — latency
2band breaches
—model half · needs provider URL
The kit measures 60 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 36 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
12 measured by the latest run48 need the model half
Metric
Owner
Role
Why this one
TNR
substring grader
alarm
the direction that ships a regression
pass rate
the judge
watch
moves when the model or the prompt moves
doc hit rate
retrieval
watch
upstream of every answer metric
cost per 1,000
the run
watch
repricing is silent and retrospective
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
40
different corpus — nothing is comparable
corpus.bytes
1,521,821
documents edited under the index
split.count
668
chunker changed — retrieval is a different system
split.size_p50
865
chunk shape changed
split.size_p95
1,769
chunk shape changed
dataset.rows
50
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.36
index rebuilt
all shared guards held between the recorded runs (rows 50, top_k 5) — which is why the history below is allowed to mean anything.
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
deterministic retrieval quality
exact match — any movement at all is real
50 rows
measured: run r002-retrieval (2026-08-02) reproduced run r001 to the digit — 0.980, 0.900, and the same five rows failing by name (L26 L28 L31 L32 L46)
retrieval latency
wider than 15%
50 rows, one laptop, five runs
measured: five runs on identical inputs spread 4.4% at p50 (17.39-18.16 ms) and 13.4% at p95 (18.05-20.47 ms). The band was 'wider than 1%' when only two runs existed; runs r003-flash and r003-pro both breached it, and re-deriving from five runs is what the standard requires. THE DRIVER IS MACHINE LOAD, NOT THE PIPELINE -- the widest reading was taken while a deploy was running on the same laptop, and every deterministic retrieval metric reproduced to the digit in the same run.
grader rates — TNR is the alarm
one row of the negative class — 0.20 on the fast tier, 0.167 on the deliberating tier
5 and 6 negatives
derived: the reference set has 5 and 6 negatives on 50 rows, so one row flipping moves TNR by 0.20 / 0.167; any finer band alarms on noise
judged rates with no second measurement
not yet known
—
not yet measured twice: judge.pass_rate has never been written to a run record, so same-input drift for it has no measurement behind it. This group held model.latency_p50_ms and model.output_tokens_total under the basis 'the model half has run once' until 2026-08-03 — that was true on 2026-08-01 and false from the moment r003 landed on 2026-08-02, and both metrics are derived below instead.
counted model tokens — deterministic input
exact match — any movement at all is real
50 rows
measured: four runs, two per tier. The prompt is assembled from the same 50 questions and the same top_k 5 passages, so the input token count reproduced TO THE DIGIT within each tier — 74,603 twice on the fast tier and 70,653 twice on the deliberating tier. Same signature that earned retrieval.doc_hit_rate its band. The two tiers differ from each other, which is why the reference is same-tier and never cross-tier.
counted model tokens — generated output
wider than 3%
50 rows
measured: four runs, two per tier. Output length is NOT deterministic even on identical input — 7,727 to 7,583 on the fast tier (1.86%) and 14,881 to 14,585 on the deliberating tier (1.99%). The band sits just above the worst observed same-input drift, so anything wider than it is real. Input tokens reproduce exactly and output tokens do not; publishing one band for both would have been wrong in one direction.
model latency — typical
wider than 6%
50 rows
measured: four runs, two per tier. p50 moved 2,023.03 to 1,944.95 ms on the fast tier (3.86%) and 4,642.61 to 4,835.03 ms on the deliberating tier (4.14%) on identical inputs. The band sits above the worse of the two. Comparing ACROSS tiers is meaningless here — the tiers are 2.4x apart — which is what made this look underivable until the reference was scoped to the same model.
model latency — tail
wider than 35%
50 rows
measured: four runs, two per tier, and the tail is far noisier than the median. p95 moved 3,735.91 to 3,902.42 ms on the fast tier (4.46%) but 8,472.65 to 10,979.50 ms on the deliberating tier — 29.59% on identical inputs. The band is set by the worse tier, which makes it wide and nearly useless as an alarm. THAT IS THE FINDING: a p95 that moves 30% run to run cannot be held to anything tighter, and a narrower band would fire on noise and be switched off within a week.
judged rates — the spread is compound
wider than 10%
50 rows
measured: four runs, two per tier. Grounding moved 0.74 to 0.80 on the fast tier (8.11%) and 0.78 to 0.76 on the deliberating tier (2.56%); judge accuracy moved 0.90 to 0.88 and 0.88 to 0.90 (2.22% and 2.27%). The band is set by the worst of the four. ⚠︎ THIS SPREAD IS COMPOUND AND CANNOT BE ATTRIBUTED: a judged rate moves when the model answers differently AND when the grader judges differently, and nothing here separates the two. So the band says only that a wider move is real — it does not say what moved, and it must never be read as a measurement of the model alone.
counted in rows, where a percentage would mislead
5 to 8 wrong of 50 — read in rows, not rates
50 rows
measured: four runs, two per tier — 8 then 5 on the fast tier, 6 then 7 on the deliberating tier. ⚠︎ NO PERCENTAGE BAND IS PUBLISHED FOR THIS ONE ON PURPOSE. Three rows out of a count of eight is 37.5%, and a band of 'wider than 45%' would look like a tolerance while actually meaning 'about three rows' — precision the denominator cannot carry. The verdict machinery compares percentages, so this band deliberately does not parse and the metric renders with no verdict. An honest blank beats a confident wrong number, which is the same rule the not-yet-known group follows.
HistoryRun history
5 recorded runs. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-flash 2026-08-01
r001-pro 2026-08-01
r002-retrieval 2026-08-02
r003-flash 2026-08-02
r003-pro 2026-08-02
Verdict
answer in the prompt
0.900
0.900
0.900
0.900
0.900
exact
doc in top-5
0.980
0.980
0.980
0.980
0.980
exact
retrieval latency p50 ms
17.52
17.39
17.61
18.16
17.81
moved +2.4%
retrieval latency p95 ms
18.05
18.49
18.20
20.47
18.50
noise +0.1%
rows failing on ranking
5
5
5
5
5
exact
judged correct
0.900
0.880
—
0.880
0.900
moved +2.3%
rows graded wrong
8
6
—
5
7
moved +16.7%
grounding rate
0.74
0.78
—
0.80
0.76
moved -2.6%
input tokens, whole run
74603
70653
—
74603
70653
exact
model latency p50 ms
2023.03
4642.61
—
1944.95
4835.03
moved +4.1%
model latency p95 ms
3735.91
8472.65
—
3902.42
10979.50
moved +29.6%
output tokens, whole run
7727
14881
—
7583
14585
moved -2.0%
What two runs bought
The deterministic metrics reproduced exactly — the same five rows fail, by name (L26 L28 L31 L32 L46) — so their band is exact-match: any movement at all is real. Latency was the only thing that drifted on identical inputs, by at most 2.5%, so its band must sit wider than that. Both bands above are those two facts.
DeviationsWhat deviated
2 breaches across 5 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
top_k down
cost down · doc hit rate down · answer-in-context down → pass rate down
measured
retrieved passages are 1,761 of 1,911 input tokens (run r001), and 4 of the balanced-set probe's 5 negatives were exactly the rows where retrieval failed
chunk size up
tokens per query up · cost up · retrieval precision down
reasoning
follows from the prompt's structure — passages dominate the token bill — but no run has varied chunk size, so it is stated as reasoning
model swap
output length up → cost up · latency p95 up · TNR moves, TPR barely
measured
run r001 across two tiers: 155 vs 298 output tokens, p95 3,736 vs 8,473 ms, substring TNR 1.000 vs 0.833 while TPR moved 0.822 to 0.864
prompt edit
grounding rate moves · every recorded verdict is invalidated
reasoning
by construction of the instrument — the recorded answers were produced by the old prompt, so re-grading them measures the old system
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
deterministic retrieval quality
nothing automatic — stated, not hidden
retrieval latency
nothing automatic — stated, not hidden
grader rates — TNR is the alarm
release gate blocks the deploy — evals/run.py
counted model tokens — deterministic input
nothing automatic — stated, not hidden
counted model tokens — generated output
nothing automatic — stated, not hidden
model latency — typical
nothing automatic — stated, not hidden
model latency — tail
nothing automatic — stated, not hidden
judged rates — the spread is compound
nothing automatic — stated, not hidden
counted in rows, where a percentage would mislead
nothing automatic — stated, not hidden
NextThe three you would add first
Treat retrieved text as data, not instructionThe measured failure. Delimit passages and state that text inside them is never an instruction — it costs nothing and addresses the one attack that worked.
A groundedness check on the way outThe faithfulness judge already exists in this kit. Running it per answer rather than per eval turns it from a measurement into a control.
A citation gateThe prompt asks for [n] markers. Nothing checks. Rejecting an uncited answer is a few lines and it is deterministic and free. OBSERVED 2026-08-03 on the live demo: asked something the corpus does not cover, the model refused correctly — "the passages do not contain any weather information" — and still cited [1][2][3][4][5]. A refusal that cites every passage it was given is the same missing check seen from the other side, and it is why the fix is a gate on the way out rather than another sentence in the prompt: the prompt here is the measured system, and editing it would make the demo a different one from every published figure.
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number. A full re-run costs about $0.21 per 1,000 rows on the published Flash-Lite card.
What this cannot tell you
Two runs is barely a history, and only the retrieval half has run twice. The model-half bands are published as not yet known rather than guessed.
This monitors THE EVAL, not a live system — there is no production traffic behind any of these numbers.
The negative class is 5 and 6 rows, so no rate band tighter than one row is honest, and the page says the denominator beside every rate.
Latency drift was measured on one machine on one day; a different machine re-baselines it.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
This kit has no framework and one dependency, pypdf. That was a choice, and the point of it is that you can read the prompt.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
extraction
src/extract.py
document loaders
the biggest genuine saving — a loader per format
chunking
src/chunk.py
text splitters
same idea, and the ceiling is the thing that matters
index and retrieval
src/retrieve.py
vector stores and retrievers
where a framework earns most: swapping keyword for embedding is a line
prompt assembly
src/prompt.py
prompt templates
44 lines here, and it is the file this kit most wants you to read
the model
src/adapters/__init__.py
chat model wrappers
a real saving, and a real abstraction cost
output
src/app.py
output parsers
nothing to parse here — the answer is prose
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing in this pipeline loops, branches or retries: one pass, seven seams. A graph earns its place when a cycle appears — re-retrieve on a bad answer, route by question type, call a tool. On this flow it would be a straight line with extra vocabulary.
The other sideWhat a framework costs you
A dependency tree, and a version treadmill attached to it.
An abstraction between you and the prompt, on a page whose whole argument is that you should see the prompt.
Defaults you did not choose — a chunk size, a template, a retry policy.
What we could NOT verify
No port was built. This maps the seams the kit already publishes; it does not prove a framework port reproduces the numbers.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-pro on the deliberating tier, 2026-08-01. This kit records telemetry measured per run, split by stage — 6 of the 6 readings on this axis, 6 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Retrieval, median
17 ms
wider than 15%
nothing automatic — stated, not hidden
Retrieval, p95
18 ms
wider than 15%
nothing automatic — stated, not hidden
Model, median
4,643 ms
wider than 6%
nothing automatic — stated, not hidden
Model, p95
8,473 ms
wider than 35%
nothing automatic — stated, not hidden
Input tokens
70,653
exact match — any movement at all is real
nothing automatic — stated, not hidden
Output tokens
14,881
wider than 3%
nothing automatic — stated, not hidden
No movement column. Not one of the 4 earlier runs on record is comparable with this one — a different model tier, or a changed guard — and differencing across that measures two systems rather than one. Said once here rather than as a dash on every row.
Where the time goesRetrieval against the model
model4,643 ms · 99.6%
retrieval17 ms · 0.4%
The only stage split anything in this estate records. It is two stages, not a span tree — the pipeline is not instrumented per step, and nothing here should be read as though it were.
DriftMedian call, run over run
r001-flash2,023 ms
r001-pro4,643 ms
r003-flash1,945 ms
r003-pro4,835 ms
Hatched bars are runs that are not comparable with the latest — a different model tier or a changed guard. They are drawn because hiding them would make the history look shorter than it is, and greyed because differencing across them would be measuring two systems.
1 run not plotted. r002-retrieval recorded no call latency at all — a rules-only floor run makes no model call, and the 0 its record stores is an absence, not a zero-millisecond answer. Plotting it would put a bar on this chart claiming the fastest run in the kit’s history.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
The split is retrieval versus model only — two stages, not the full pipeline.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: a deterministic chunk index in a file on your disk.
The machineWhat this needs
Dependencies
pypdf>=4.0
Configuration
8 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-02, across 5 committed records
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
corpus
data/corpus/ — your disk
never whole; only the retrieved passages, per call
chunk index
chunks.json — your disk, deterministic
never
labels
data/labelled.jsonl
only during the free eval, to the judge you configure
the key
.env — never committed
only inside the Authorization header
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
src/app.py line 110
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
tools/fetch_corpus.py line 52
https://www.sqlite.org/%s.html
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the environment only, never written into the repo, never requested from a reader on any surface.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
668 chunks score in 17ms; the scan is linear (lenses.Architecture.breaks_at_scale, cold-clone verified 2026-07-31)
pgvector on the Postgres you already run; a dedicated store past that. The seam is src/retrieve.py — keyword and embedding sit behind one call, and dispatch is on what the index HAS, not what .env asks for
published accuracy does not transfer to a new retriever — your labels stay valid, and the eval re-runs free
chunking
paragraph-greedy to a hard 1,800-character ceiling (src/chunk.py)
chunk ids change — the index rebuilds and citations re-anchor
model
one HTTP completion call behind src/adapters/__init__.py — any OpenAI-compatible endpoint or Anthropic; a local server is the zero-egress path
—
a hosted provider for quality, a local Ollama / vLLM / LM Studio server for data that cannot leave — the .env decides, not the code
published quality numbers are per-model — the eval harness re-runs free on yours
corpus refresh
delete, replace, python -m src.index — a whole-index rebuild, no increments
0.36s at 40 documents (lenses.Data.index.build_seconds)
a scheduled rebuild, until the rebuild itself is the bottleneck — which no run has measured past this corpus
nothing — the index is deterministic, so a rebuild that changes it is a real diff, not noise
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
the right document, the wrong passage — the table of contents outranks the content
a heading split from its answer at the 1,800-character ceiling; four of the five recorded retrieval failures are this class — L26, L28, L31 and L46, whose document WAS retrieved. L32 is not: its document (queryplanner) never appeared in the top 5 at all, so no split inside it can explain the miss
re-cut at the splitter seam (src/chunk.py), then re-run retrieval (lenses.Data.breaks_on — the five recorded failures)
a document never appears in any answer
a scanned or image-only PDF — pypdf extracts no text and there is no OCR step
check extraction output per document before blaming retrieval (lenses.Data.breaks_on — recorded against this corpus)
a file describing the corpus answers questions about the corpus
SOURCES.md was silently indexed as a 41st document until 2026-08-01
keep metadata files out of data/corpus/ — src/index.py names the exclusion (lenses.Data.breaks_on, fifth entry)
Concurrency and GPU sizing — no run produced them, so they are absent rather than estimated. Provider-side retention, training use and log residency — provider-dependent, a third state. And the 17ms retrieval scan past this corpus: linear is the measured mechanism, not a measurement at your scale.
The corpus licence, from the Data lens: Public domain. SQLite's dedication names documentation as well as code, verified against sqlite.org/copyright.html on 2026-07-31. No attribution clause, no share-alike, no notice file. Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineExact substring
Does the answer contain a required fragment, verbatim?
$0.00per 1,000 questions
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/run.py, in-process, no key. The fragment comes from answer_contains in data/labelled.jsonl.
Every grader on these pages scored the same 100 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The row
L21
The question
How does WAL mode change the blocking relationship between readers and writers?
The reference answer
Readers do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
What the system answered
In WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
The required fragment
proceed concurrently
What retrieval returned
all five passages came from the wal document
The answer is right. It says “run at the same time” where the label says “proceed concurrently”, so the substring grader fails it and every other grader passes it.
Grader
Verdict
Why
Exact substring
fail
the fragment is absent
Token F1
pass
0.273, threshold 0.05
ROUGE-L
pass
0.273, threshold 0.05
Reference recall
pass
0.500, threshold 0.05
LLM judge, zero-shot
pass
“Both state that readers and writers do not block each other and can run concurrently.”
LLM judge, few-shot
pass
unchanged from zero-shot
Faithfulness
pass
supported by the retrieved passages
The formulaWhat it computes
fragment.lower() in answer.lower()
IterationsFour rulers, four wrong numbers
#
What changed
What it found
1
the first labels used whatever phrase appeared in the document
a table-of-contents entry was chosen as a fragment; it is genuinely in the document and answers nothing
2
MAX_FRAGMENT_WORDS = 4
measured over 100 graded answers: 80% correct at 1-2 word fragments, 33% at 7+, on the same answers. The ruler was scoring label length
3
check 6 forbids polar questions
“Yes.” is a correct answer to “does WAL mode affect concurrency?” and contains no fragment, so the grader rewarded padding
4
abandoned as the correctness grader
paraphrase rigidity. Both models sat near 86% true accuracy while this reported 78% and 74% — and reported them in the wrong ORDER
The tell was the ranking inverting, not the value moving. Noise moves a number; an artifact reorders the comparison the page exists to make.
The analysisWhat it actually did
Model
Result
the fast tier
agreed with the judge 84.0% · -6.0 against the null grader · 0 false pass, 8 false fail
the deliberating tier
agreed with the judge 86.0% · -2.0 against the null grader · 1 false pass, 6 false fail
In operationWhat to monitor
Reference standard: adjudicated truth, not another grader — all 100 answers were read and given a verdict, including the 85 the two graders already agreed on, so these are accuracy and not agreement. The adjudicator is itself a model, so this is model-confirmed and not human-confirmed.
Model
TPR
TNR — the alarm metric
Precision
Pass rate
fp/fn
one row moves TNR by
the fast tier
0.822
1.000
1.000
74%
0/8
±0.20
the deliberating tier
0.864
0.833
0.974
78%
1/6
±0.17
both tiers — all 100 answers
0.843
0.909
0.987
76%
1/14
±0.09
Watch these
TNR — the alarm metric
TPR
precision
pass rate
Alarm on
TNR — it is the direction that ships a regression. A grader that fails good answers wastes review time; one that passes bad answers is why the bad answer reaches a user.
How tight can the band be? Over all 100 answers this check is right 85 times, and its error is almost entirely in one direction: 14 answers it failed that the adjudication scored correct, against 1 it passed that was wrong. The negative class is 11 rows of 100 — wider than the 5 and 6 a single model gives, but one row still moves TNR by 0.09. The band stays one row wide and the denominator is printed beside every rate.
Cadence: Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number.
The decisionWhen to reach for it
Use it
As a GROUNDING signal: did the expected content reach the answer verbatim.
Do not use it
As a correctness score. It fails in one direction only, and a biased grader moves the headline rather than widening the error bar.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineToken F1
How much of the reference wording overlaps the answer, in both directions?
$0.00per 1,000 questions
nodata leaves your network
yessame answer every time
MethodHow the test was run
Offline over the committed result files. Content words only; stopwords dropped.
Every grader on these pages scored the same 100 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The row
L21
The question
How does WAL mode change the blocking relationship between readers and writers?
The reference answer
Readers do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
What the system answered
In WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
The required fragment
proceed concurrently
What retrieval returned
all five passages came from the wal document
The answer is right. It says “run at the same time” where the label says “proceed concurrently”, so the substring grader fails it and every other grader passes it.
Grader
Verdict
Why
Exact substring
fail
the fragment is absent
Token F1
pass
0.273, threshold 0.05
ROUGE-L
pass
0.273, threshold 0.05
Reference recall
pass
0.500, threshold 0.05
LLM judge, zero-shot
pass
“Both state that readers and writers do not block each other and can run concurrently.”
LLM judge, few-shot
pass
unchanged from zero-shot
Faithfulness
pass
supported by the retrieved passages
The formulaWhat it computes
harmonic mean of precision and recall over the multiset of content-word tokens
A continuous score needs a decision threshold to become pass/fail, and the threshold is a knob someone picks. This one was swept 0.05 to 0.95 in steps of 0.05 and the best-agreeing value was ≥ 0.05. The sweep is published with the result, because a single tuned number shown alone is the same trick as quoting a test score you tuned on.
The analysisWhat it actually did
Model
Result
the fast tier
agreed with the judge 96.0% · +6.0 against the null grader · 2 false pass, 0 false fail
the deliberating tier
agreed with the judge 92.0% · +4.0 against the null grader · 4 false pass, 0 false fail
In operationWhat to monitor
Reference standard: the LLM judge, zero-shot — so these are agreement split by DIRECTION, not accuracy against truth.
Model
TPR
TNR — the alarm metric
Precision
Pass rate
fp/fn
one row moves TNR by
the fast tier
1.000
0.600
0.957
94%
2/0
±0.20
the deliberating tier
1.000
0.333
0.917
96%
4/0
±0.17
Watch these
TNR — the alarm metric
TPR
precision
pass rate
Alarm on
TNR — it is the direction that ships a regression. A grader that fails good answers wastes review time; one that passes bad answers is why the bad answer reaches a user.
How tight can the band be? The reference has only 5 and 6 negatives on 50 rows, so ONE row flipping moves TNR by 0.20 and 0.17. Any alarm band finer than one row is alarming on noise — which is why the band here is one row wide and the denominator is printed beside it.
Cadence: Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number.
The decisionWhen to reach for it
Use it
A free, private first pass, and the closest cheap approximation of the judge here.
Do not use it
When the right answer can use entirely different words from the reference. Overlap cannot see a synonym.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineReference recall
How much of the reference answer's content made it into the answer?
$0.00per 1,000 questions
nodata leaves your network
yessame answer every time
MethodHow the test was run
Offline over the committed result files.
Every grader on these pages scored the same 100 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The row
L21
The question
How does WAL mode change the blocking relationship between readers and writers?
The reference answer
Readers do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
What the system answered
In WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
The required fragment
proceed concurrently
What retrieval returned
all five passages came from the wal document
The answer is right. It says “run at the same time” where the label says “proceed concurrently”, so the substring grader fails it and every other grader passes it.
Grader
Verdict
Why
Exact substring
fail
the fragment is absent
Token F1
pass
0.273, threshold 0.05
ROUGE-L
pass
0.273, threshold 0.05
Reference recall
pass
0.500, threshold 0.05
LLM judge, zero-shot
pass
“Both state that readers and writers do not block each other and can run concurrently.”
LLM judge, few-shot
pass
unchanged from zero-shot
Faithfulness
pass
supported by the retrieved passages
The formulaWhat it computes
share of the reference's content words that appear in the answer
A continuous score needs a decision threshold to become pass/fail, and the threshold is a knob someone picks. This one was swept 0.05 to 0.95 in steps of 0.05 and the best-agreeing value was ≥ 0.05. The sweep is published with the result, because a single tuned number shown alone is the same trick as quoting a test score you tuned on.
The analysisWhat it actually did
Model
Result
the fast tier
agreed with the judge 96.0% · +6.0 against the null grader · 2 false pass, 0 false fail
the deliberating tier
agreed with the judge 92.0% · +4.0 against the null grader · 4 false pass, 0 false fail
In operationWhat to monitor
Reference standard: the LLM judge, zero-shot — so these are agreement split by DIRECTION, not accuracy against truth.
Model
TPR
TNR — the alarm metric
Precision
Pass rate
fp/fn
one row moves TNR by
the fast tier
1.000
0.600
0.957
94%
2/0
±0.20
the deliberating tier
1.000
0.333
0.917
96%
4/0
±0.17
Watch these
TNR — the alarm metric
TPR
precision
pass rate
Alarm on
TNR — it is the direction that ships a regression. A grader that fails good answers wastes review time; one that passes bad answers is why the bad answer reaches a user.
How tight can the band be? The reference has only 5 and 6 negatives on 50 rows, so ONE row flipping moves TNR by 0.20 and 0.17. Any alarm band finer than one row is alarming on noise — which is why the band here is one row wide and the denominator is printed beside it.
Cadence: Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number.
The decisionWhen to reach for it
Use it
When completeness matters more than concision.
Do not use it
It cannot punish a long answer at all — padding scores well. That is the opposite bias to the substring grader and just as real.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineROUGE-L
Does the answer follow the reference's wording in the same ORDER?
$0.00per 1,000 questions
nodata leaves your network
yessame answer every time
MethodHow the test was run
Offline over the committed result files.
Every grader on these pages scored the same 100 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The row
L21
The question
How does WAL mode change the blocking relationship between readers and writers?
The reference answer
Readers do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
What the system answered
In WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
The required fragment
proceed concurrently
What retrieval returned
all five passages came from the wal document
The answer is right. It says “run at the same time” where the label says “proceed concurrently”, so the substring grader fails it and every other grader passes it.
Grader
Verdict
Why
Exact substring
fail
the fragment is absent
Token F1
pass
0.273, threshold 0.05
ROUGE-L
pass
0.273, threshold 0.05
Reference recall
pass
0.500, threshold 0.05
LLM judge, zero-shot
pass
“Both state that readers and writers do not block each other and can run concurrently.”
LLM judge, few-shot
pass
unchanged from zero-shot
Faithfulness
pass
supported by the retrieved passages
The formulaWhat it computes
F-measure over the longest common subsequence of content words
A continuous score needs a decision threshold to become pass/fail, and the threshold is a knob someone picks. This one was swept 0.05 to 0.95 in steps of 0.05 and the best-agreeing value was ≥ 0.05. The sweep is published with the result, because a single tuned number shown alone is the same trick as quoting a test score you tuned on.
The analysisWhat it actually did
Model
Result
the fast tier
agreed with the judge 94.0% · +4.0 against the null grader · 2 false pass, 1 false fail
the deliberating tier
agreed with the judge 90.0% · +2.0 against the null grader · 4 false pass, 1 false fail
In operationWhat to monitor
Reference standard: the LLM judge, zero-shot — so these are agreement split by DIRECTION, not accuracy against truth.
Model
TPR
TNR — the alarm metric
Precision
Pass rate
fp/fn
one row moves TNR by
the fast tier
0.978
0.600
0.957
92%
2/1
±0.20
the deliberating tier
0.977
0.333
0.915
94%
4/1
±0.17
Watch these
TNR — the alarm metric
TPR
precision
pass rate
Alarm on
TNR — it is the direction that ships a regression. A grader that fails good answers wastes review time; one that passes bad answers is why the bad answer reaches a user.
How tight can the band be? The reference has only 5 and 6 negatives on 50 rows, so ONE row flipping moves TNR by 0.20 and 0.17. Any alarm band finer than one row is alarming on noise — which is why the band here is one row wide and the denominator is printed beside it.
Cadence: Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number.
The decisionWhen to reach for it
Use it
When sequence carries meaning — steps, ordered procedures.
Do not use it
For a one-clause factual answer, order is noise and this is Token F1 with a penalty attached.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineLLM judge, zero-shot
Is the answer correct, judged on meaning rather than wording?
$0.36per 1,000 questions
yesdata leaves your network
nosame answer every time
MethodHow the test was run
evals/judge.py against a completed run. The judge NEVER sees the fragment — showing it would rebuild the instrument being replaced.
Every grader on these pages scored the same 100 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The row
L21
The question
How does WAL mode change the blocking relationship between readers and writers?
The reference answer
Readers do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
What the system answered
In WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
The required fragment
proceed concurrently
What retrieval returned
all five passages came from the wal document
The answer is right. It says “run at the same time” where the label says “proceed concurrently”, so the substring grader fails it and every other grader passes it.
Grader
Verdict
Why
Exact substring
fail
the fragment is absent
Token F1
pass
0.273, threshold 0.05
ROUGE-L
pass
0.273, threshold 0.05
Reference recall
pass
0.500, threshold 0.05
LLM judge, zero-shot
pass
“Both state that readers and writers do not block each other and can run concurrently.”
LLM judge, few-shot
pass
unchanged from zero-shot
Faithfulness
pass
supported by the retrieved passages
The promptWhat the judge is asked
6 of 6 — four fluent wrong answers rejected, two correct paraphrases accepted
The analysisWhat it actually did
Model
Result
the fast tier
scored 90.0%
the deliberating tier
scored 88.0%
54 calls for 50 rows and 56 for 50; the extra calls are ceiling escalations, because a token cap covers thinking and not just output
In operationWhat to monitor
Reference standard: none, because this grader IS the reference standard. Every other correctness grader on these pages is scored against its verdicts, which is exactly why its own cannot be scored from anything in this table.
These rates are UNKNOWN, on purpose
This grader's TPR and TNR are UNKNOWN and are printed as unknown. It is the reference standard the others are measured against, and a grader cannot be scored against itself — a rate here would be circular and would read as evidence. All 100 answers have since been adjudicated and this grader matched the adjudication on every one, but the adjudicator is another model built the same way, so that is validated against a model adjudication and is not human-confirmed. It is deliberately not published as a rate. What would fill these in is a person reading the rows.
Watch these
pass rate against its own baseline
unreadable-verdict count
calls per row (ceiling escalations)
spend per 1,000 rows
Alarm on
TNR — it is the direction that ships a regression. A grader that fails good answers wastes review time; one that passes bad answers is why the bad answer reaches a user.
How tight can the band be? The reference has only 5 and 6 negatives on 50 rows, so ONE row flipping moves TNR by 0.20 and 0.17. Any alarm band finer than one row is alarming on noise — which is why the band here is one row wide and the denominator is printed beside it.
Cadence: Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number.
The decisionWhen to reach for it
Use it
When the criterion needs interpretation and you accept the bill and the exposure.
Do not use it
When a rule could decide it. A judge is the last resort, not the first reach.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineLLM judge, few-shot
The same question, with two worked examples in the prompt.
$0.36per 1,000 questions
yesdata leaves your network
nosame answer every time
MethodHow the test was run
Same 100 recorded answers, same judge model, two examples added: a terse polar answer that must pass, and a fluent wrong one that must fail.
Every grader on these pages scored the same 100 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The row
L21
The question
How does WAL mode change the blocking relationship between readers and writers?
The reference answer
Readers do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
What the system answered
In WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
The required fragment
proceed concurrently
What retrieval returned
all five passages came from the wal document
The answer is right. It says “run at the same time” where the label says “proceed concurrently”, so the substring grader fails it and every other grader passes it.
Grader
Verdict
Why
Exact substring
fail
the fragment is absent
Token F1
pass
0.273, threshold 0.05
ROUGE-L
pass
0.273, threshold 0.05
Reference recall
pass
0.500, threshold 0.05
LLM judge, zero-shot
pass
“Both state that readers and writers do not block each other and can run concurrently.”
LLM judge, few-shot
pass
unchanged from zero-shot
Faithfulness
pass
supported by the retrieved passages
The promptWhat the judge is asked
The zero-shot criterion, unchanged, with two worked examples appended: a terse polar answer that must pass, and a fluent wrong one that must fail.
The analysisWhat it actually did
Model
Result
the fast tier
scored 90.0% · agreed with the judge 100.0% · 0 false pass, 0 false fail
the deliberating tier
scored 88.0% · agreed with the judge 100.0% · 0 false pass, 0 false fail
The result
ZERO verdicts moved, on either model.
In operationWhat to monitor
Reference standard: the LLM judge, zero-shot — so these are agreement split by DIRECTION, not accuracy against truth.
These rates are UNKNOWN, on purpose
This grader's TPR and TNR against truth are UNKNOWN and are printed as unknown. It is NOT the reference standard — the zero-shot judge is — and against that reference it moved zero verdicts on either model. Perfect agreement with the reference is agreement, not accuracy: on any row the reference has wrong, this grader is wrong with it and the agreement figure cannot show it.
Watch these
pass rate against its own baseline
unreadable-verdict count
calls per row (ceiling escalations)
spend per 1,000 rows
Alarm on
TNR — it is the direction that ships a regression. A grader that fails good answers wastes review time; one that passes bad answers is why the bad answer reaches a user.
How tight can the band be? The reference has only 5 and 6 negatives on 50 rows, so ONE row flipping moves TNR by 0.20 and 0.17. Any alarm band finer than one row is alarming on noise — which is why the band here is one row wide and the denominator is printed beside it.
Cadence: Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number.
The decisionWhen to reach for it
Use it
When the criterion is genuinely ambiguous and the examples pin down a boundary.
Do not use it
Here. The examples bought nothing but tokens, and that is a measurement rather than an opinion.
Answer customer questions from your product manuals
PresenterOpens the private repo. Visible to admins only.
In one lineFaithfulness
Does the answer follow from the passages it was given — regardless of whether it is correct?
$0.36per 1,000 questions
yesdata leaves your network
nosame answer every time
MethodHow the test was run
The retrieved passages were RE-DERIVED (the result files record chunk ids, not text) and the derivation proved: regenerated chunk ids matched the recorded ones on 50 of 50 rows, both models, before anything was judged.
Every grader on these pages scored the same 100 already-recorded answers. Nothing was re-generated, so this compares rulers and not models.
The inputOne real row, seen by every grader
The row
L21
The question
How does WAL mode change the blocking relationship between readers and writers?
The reference answer
Readers do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
What the system answered
In WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
The required fragment
proceed concurrently
What retrieval returned
all five passages came from the wal document
The answer is right. It says “run at the same time” where the label says “proceed concurrently”, so the substring grader fails it and every other grader passes it.
Grader
Verdict
Why
Exact substring
fail
the fragment is absent
Token F1
pass
0.273, threshold 0.05
ROUGE-L
pass
0.273, threshold 0.05
Reference recall
pass
0.500, threshold 0.05
LLM judge, zero-shot
pass
“Both state that readers and writers do not block each other and can run concurrently.”
LLM judge, few-shot
pass
unchanged from zero-shot
Faithfulness
pass
supported by the retrieved passages
The promptWhat the judge is asked
4 of 4 — a fabricated number, a flat contradiction and a claim that is TRUE but absent from the passages were all rejected; a genuine paraphrase passed
The analysisWhat it actually did
Model
Result
the fast tier
scored 100.0%
the deliberating tier
scored 98.0%
The result
100% and 98% is a real result, not a broken gate — and it was red-proved before being believed.
In operationWhat to monitor
Reference standard: none in this table, and that is not an omission. Every other grader here answers "is the answer correct" and is scored against the zero-shot judge. This one asks whether the answer follows from its passages, so agreeing or disagreeing with a correctness judge would mean nothing in either direction. It was checked by a 4-case red-proof instead.
These rates are UNKNOWN, on purpose
This grader's TPR and TNR are UNKNOWN and are printed as unknown, for a different reason from the zero-shot judge's. That one cannot be scored against itself. This one has nothing here to be scored against at all — it decides a different question, so there is no reference standard in this table whose verdicts its own could be compared with.
Watch these
pass rate against its own baseline
unreadable-verdict count
calls per row (ceiling escalations)
spend per 1,000 rows
Alarm on
TNR — it is the direction that ships a regression. A grader that fails good answers wastes review time; one that passes bad answers is why the bad answer reaches a user.
How tight can the band be? It flagged nothing at all on the fast tier and one answer on the deliberating tier. A negative class with no members cannot produce a rate, and one with a single member swings the whole way if that row flips — so there is no alarm band here worth drawing, and the counts are printed instead of a percentage. The 4-case red-proof is the evidence that it can say no; 100 percent is not.
Cadence: Grade a sample on every release, and the whole set when the model, the prompt, the chunker or top_k changes — those four are the seams that move the number.
The decisionWhen to reach for it
Use it
To catch hallucination, which a correctness grader cannot see.
Do not use it
As a correctness score. An answer can be perfectly faithful to passages that were themselves the wrong ones.
Scoring the graders themselves, against rows a second reader adjudicated — not scoring the system.
PresenterOpens the private repo. Visible to admins only.
Read this firstWhat this measures, and what it does not
Error analysis reads the system’s failures. Adjudication settles a fight between two graders. This is the third thing — using those settled rows to score the graders. The system is not on trial here; the rulers are.
Now measured
Every one of the 50 rows was adjudicated on both executions — the 15 the graders fought over and the 85 they agreed on. Agreement is not correctness, so the agreeing rows had to be read too. Nothing below is computed from a selected sample.
100
Answers adjudicated
all 50 rows, both executions — not a sample
100%
The LLM judge agreed
with the adjudicated answer
85%
The substring check agreed
85 of 100 — it is not worthless, it just loses arguments
—
A person confirmed this
still empty, and that is the point
Each row was read twice by a second model — once plainly, once by a reader told to hunt for a reason to fail it. The two passes agreed on 15 of 15. That is a useful signal and it is not human confirmation. These numbers have since been written onto the Evals pages: the substring check now publishes 85.0% against adjudicated truth, and the judge’s own rates stay UNKNOWN, because a grader cannot be scored against itself and the adjudicator here is another model. ingest_adjudication.py was not the path and is superseded — its whole scope is the 15 disputed rows.
ScoresHow each grader scored
The LLM judge
Agreed with the adjudicated answer100 of 100
100%
True positive rate — found the correct answers89 of 89
100%
True negative rate — caught the wrong ones11 of 11
100%
Precision — when it said correct, it was89 of 89
100%
Exact substring
Agreed with the adjudicated answer85 of 100
85%
True positive rate — found the correct answers75 of 89
84%
True negative rate — caught the wrong ones10 of 11
91%
Precision — when it said correct, it was75 of 76
99%
The LLM judge
Agreed with the adjudicated answer50 of 50
100%
True positive rate — found the correct answers45 of 45
100%
True negative rate — caught the wrong ones5 of 5
100%
Precision — when it said correct, it was45 of 45
100%
Exact substring
Agreed with the adjudicated answer42 of 50
84%
True positive rate — found the correct answers37 of 45
82%
True negative rate — caught the wrong ones5 of 5
100%
Precision — when it said correct, it was37 of 37
100%
The LLM judge
Agreed with the adjudicated answer50 of 50
100%
True positive rate — found the correct answers44 of 44
100%
True negative rate — caught the wrong ones6 of 6
100%
Precision — when it said correct, it was44 of 44
100%
Exact substring
Agreed with the adjudicated answer43 of 50
86%
True positive rate — found the correct answers38 of 44
86%
True negative rate — caught the wrong ones5 of 6
83%
Precision — when it said correct, it was38 of 39
97%
Every bar carries its denominator. Read the substring row against the earlier version of this page: on the 15 disputed rows it scored 0 of 15 and looked worthless. Over all fifty it is 85% accurate. It is not a bad grader — it is a grader that loses every argument it picks, which is a different and much more useful thing to know.
The four outcomesCounts, before any rate
Grader
Said correct, was correct
Said correct, was wrong
Said wrong, was wrong
Said wrong, was correct
The LLM judge
89
0
11
0
Exact substring
75
1
10
14
Read the second column first. It is the expensive one: a grader that says correct about a wrong answer publishes a number and nothing looks red. A grader that says wrong about a correct answer only wastes a review. The substring check has the only entry in that column, and it is row 11.
Not yet knownWhy the judge scores 100%
Read this before quoting it
The judge agrees with the adjudication on every row — and the adjudicator is also a model. Two systems built the same way can share a blind spot, so a perfect score here is evidence they agree, not proof they are right.
The substring check does not have this problem: it is deterministic and shares nothing with either. Its 85% is the more trustworthy figure on this page. The one row that would settle the judge’s score is a person reading a sample — which is why that card stays dashed and ingest_adjudication.py is still unrun.
If you grade another waySix methods, and what each one needs
This kit used an LLM judge. That is one method, not the method. If you grade by hand, by rubric, by panel or by a deterministic script, your inputs, your outputs, the way you validate the method, and the dashboard you need are all different. Open any row.
Cost is denominated in the method’s own unit. A human reviewer costs minutes, not tokens — a manual method reported as $0.00 is not cheap, it is unpriced.
Other methods
This kit graded with an LLM as judge, and everything on this page is that method being measured. If you grade by hand, by rubric, by panel, by a deterministic script or by routing between them, the dashboard you need is a different shape — all 6 are laid out, with their inputs, outputs and validation techniques, on Evals & Testing.
They are kept on a separate page on purpose. Those panels are templates — illustrative shapes for methods we did not run — and a template sitting beside a measurement is how the two get mistaken for each other.
The casesThe fifteen disputed rows, in full
Click a case for the answer given, the reference, the fragment the substring check wanted, and both readings. Nothing here is summarised away.
Answer givenIn WAL mode, readers and writers do not block each other; they can run at the same time. [1][4]
ReferenceReaders do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
Fragment wantedproceed concurrently
Judge saidBoth state that readers and writers do not block each other and can run concurrently.
Plain reading"Readers and writers do not block each other; they can run at the same time" is exactly the WAL concurrency property the reference states, just in plainer words; the missing phrase "proceed concurrently" is wording, not fact. Agrees with the LLM judge.
Skeptical reading"Readers and writers do not block each other; they can run at the same time" is exactly the WAL property the question asks about, and it makes no over-claim (it does not say multiple writers can run concurrently). Agrees with the judge; the substring check only failed because the paraphrase avoids the words "proceed concurrently".
Answer givenThe documentation recommends avoiding the AUTOINCREMENT keyword unless it is strictly needed, because it imposes extra CPU, memory, disk space, and disk I/O overhead and makes INSERTs run slower [3][4]. It is usually not needed, since a plain INTEGER PRIMARY KEY already automatically assigns unused ROWIDs, typically one more than the largest currently in use [3]. AUTOINCREMENT should only be used if the application requires the guaranteed properties that automatically chosen ROWIDs are never reused and are monotonically increasing [4].
ReferenceThat it should be avoided if not strictly needed, because it imposes extra CPU, memory, disk space and disk I/O overhead.
Fragment wantedshould be avoided
Judge saidThe candidate correctly states the recommendation to avoid AUTOINCREMENT unless strictly needed and cites the overhead reasons; adding detail does not contradict.
Plain readingIt gives the recommendation (avoid AUTOINCREMENT unless strictly needed) and the stated reason (extra CPU, memory, disk space and disk I/O), and the added detail — INTEGER PRIMARY KEY already assigns an unused ROWID, usually one more than the largest, and AUTOINCREMENT only guarantees no reuse plus monotonic increase — is also true of SQLite. Agrees with the LLM judge.
Skeptical readingIt states the recommendation (avoid unless strictly needed) and the stated reason (extra CPU, memory, disk space and disk I/O overhead); the added detail — slower INSERTs from maintaining sqlite_sequence, plain INTEGER PRIMARY KEY normally choosing largest rowid + 1, and the never-reused/monotonic guarantees — is all accurate SQLite behaviour, so nothing false is asserted. Agrees with the judge.
Answer givenA WITHOUT ROWID table can reduce disk space usage and/or improve performance, since omitting the rowid sometimes uses less disk space and runs a little faster [1][2].
ReferenceIt is an optimization that can reduce storage and processing requirements.
Fragment wantedreduce storage
Judge saidThe candidate correctly describes reductions in disk space and performance improvements, matching the reference's optimization of storage and processing requirements.
Plain reading"Can reduce disk space usage and/or improve performance" is the same claim as the reference's "reduce storage and processing requirements", and matches SQLite's own wording that WITHOUT ROWID tables can use less disk space and run a little faster. Agrees with the LLM judge.
Skeptical reading"Reduce disk space usage and/or improve performance" covers both halves of the reference's "storage and processing requirements", and hedging with "sometimes" / "a little faster" is conservative rather than wrong (the docs claim up to half the space and nearly twice the speed in some cases). Agrees with the judge.
Answer givenAn index can also reference expressions involving table columns. [4]
ReferenceAn index can be formed on expressions involving table columns, either by writing an expression in the CREATE INDEX statement or by indexing a generated column.
Fragment wantedformed on expressions
Judge saidThe candidate correctly states that an index can reference expressions involving table columns, which is the core of the reference answer.
Plain readingThe question asks what else an index can reference besides table columns, and "expressions involving table columns" is exactly right; the reference's extra detail about how those expressions are written (CREATE INDEX expression or generated column) is elaboration, not a separate fact the answer contradicts. Agrees with the LLM judge.
Skeptical readingThe question asks only what else an index can reference besides table columns, and "expressions involving table columns" is exactly that answer; it omits the reference's two mechanisms (expression in CREATE INDEX / generated column) but omission is not error, and nothing asserted is false. Agrees with the judge.
Answer givenSQLite uses the rowid as the access key for the data in the underlying B-tree storage engine, making rowid access highly optimized and fast [2].
ReferenceDirectly, via the rowid, which is the fastest lookup available.
Fragment wantedbinary search
Judge saidConveys that rowid lookup is direct and fast, consistent with reference.
Plain readingSaying SQLite uses the rowid as the access key into the underlying B-tree, making the lookup highly optimized and fast, is a factually accurate description of direct rowid lookup and asserts nothing false; the missing phrase "binary search" is a substring-check artefact, not an error. Agrees with the LLM judge.
Skeptical readingPressed hardest here, since the answer describes the storage arrangement (rowid is the access key into the B-tree) rather than narrating the lookup, and never says "binary search" — but that mechanism is factually right for SQLite, implies the direct rowid lookup the reference describes, and asserts nothing false; the required fragment is not even present in the reference answer itself. Agrees with the judge.
Answer givenAll database reads and writes occur only within a transaction. Any command that accesses the database will automatically start a transaction if one is not already in effect, and automatically started transactions are committed when the last SQL statement finishes [1].
ReferenceNo reads or writes occur except within a transaction. Any command that accesses the database will automatically start one if it is not already in progress.
Fragment wantedno reads or writes
Judge saidThe candidate correctly paraphrases the reference's key point and includes an additional accurate detail about automatic transaction commitment.
Plain readingIt states both halves of the reference — reads and writes occur only inside a transaction, and any database-accessing command auto-starts one if none is in effect — and the added claim that an automatically started transaction commits when the last statement finishes is also correct per the SQLite docs. Agrees with the LLM judge.
Skeptical readingIt states both halves of the reference (access only within a transaction; any accessing command auto-starts one), and the extra claim that auto-started transactions commit when the last statement finishes matches SQLite's transaction documentation, so the added detail does not introduce an error. Agrees with the judge.
Answer givenYou query an R*Tree index by using standard SQL expressions in the WHERE clause of a SELECT query to find entries that intersect with or are contained within a bounding-box, or by using the MATCH operator for custom regions [3].
ReferenceBy constraining the min and max columns of the R*Tree virtual table.
Fragment wantedmatch operator
Judge saidThe candidate describes querying R*Tree via WHERE clauses constraining min/max columns, which matches the reference's method, and adds correct details about bounding-box queries and MATCH operator.
Plain readingQuerying an R*Tree by putting bounding-box constraints in the WHERE clause of a SELECT is precisely what "constraining the min and max columns" means, and the MATCH operator for custom regions is a real R*Tree feature, so nothing asserted is false. I agree with the judge's verdict, though its stated reason wrongly claims the answer named the min/max columns.
Skeptical readingNothing in the answer is false: querying an R*Tree by putting bounding-box intersect/contain constraints in the WHERE clause of a SELECT is precisely what constraining the min/max columns amounts to, and MATCH with a custom geometry callback is a genuine R*Tree query path. Caveat for the record — the judge's stated reason claims the candidate 'describes querying R*Tree via WHERE clauses constraining min/max columns', and it never mentions min/max columns at all; the verdict is right, the rationale is not.
Answer givenSQLite uses anomaly testing for out-of-memory, I/O errors, crashes, and compound failure tests [1][2].
Judge saidThe candidate correctly lists out-of-memory, I/O errors, crashes, and compound failure tests, matching the reference.
Plain readingThe answer lists out-of-memory, I/O errors, crashes and compound failure tests — the same four kinds as the reference, in the same order. The substring check fails only because "testing" is not repeated after each item.
Skeptical readingThe answer names all four kinds in the reference — out-of-memory, I/O error, crash, and compound failure tests — with nothing added or contradicted; the substring check failed only because the answer factored the shared word 'testing' out of the list instead of repeating it after each item.
Answer givenIn SQLite, a database object is identified by its name and the name of the database it resides in [1].
ReferenceBy the name of the object plus the name of the database it resides in.
Fragment wantedname of the object
Judge saidThe candidate correctly states the identification method: object name plus database name.
Plain reading"Identified by its name and the name of the database it resides in" is semantically identical to the reference; the fragment name of the object misses purely because the answer writes "its name".
Skeptical reading'Identified by its name and the name of the database it resides in' is the reference answer with one pronoun substituted, and it matches SQLite's own statement that an object is identified by the object name plus the database it resides in; the substring check failed purely on 'its name' vs 'name of the object'.
Answer givenIn WAL mode, readers and writers do not block each other; they can run at the same time [1][4].
ReferenceReaders do not block writers and a writer does not block readers, so reading and writing can proceed concurrently.
Fragment wantedproceed concurrently
Judge saidThe candidate accurately paraphrases that readers and writers do not block each other and can operate concurrently.
Plain readingWAL mode does allow readers and writers to proceed at the same time without blocking one another, which is exactly what the answer says; the substring check only failed because the answer used "run at the same time" instead of the literal phrase "proceed concurrently".
Skeptical reading"Readers and writers do not block each other; they can run at the same time" is exactly the WAL behaviour the reference states, just in plainer words. Skeptical check: it never claims writers don't block other writers (which would be false), and although it states the post-WAL situation rather than explicitly contrasting it with the rollback-journal behaviour, the question's substance — the new blocking relationship — is answered correctly. Agreeing with the LLM judge; the substring grader is only penalising the phrase "proceed concurrently".
Answer givenThe provided passages do not contain the specific details about the user's responsibility regarding indexes. [3] only mentions that section 3 describes the indexes the user must create and those that should be created for efficiency, but the content of that section is not included here.
ReferenceThe user must create certain indexes in order to use foreign key constraints at all, and others should be created for those constraints to function efficiently.
Fragment wantedmust create
Judge saidThe candidate fails to answer the question, stating the details are missing, while the reference clearly outlines the user's responsibility.
Plain readingThis is a refusal — it says the passages do not contain the information — even though the question has a real answer (some indexes must be created for foreign keys to work at all, others should be created for efficiency); the substring grader scored it correct only because the quoted sentence happened to contain "must create".
Skeptical readingThis is a refusal: the system says the passages do not contain the details and merely notes that some section describes them, so it never states that the user must create certain indexes for FK constraints to work at all and others for them to be efficient. A question with a real answer answered by declining is wrong, and the substring grader scored it "correct" only because the refusal quoted the words "must create". Agreeing with the LLM judge.
Answer givenA WITHOUT ROWID table can sometimes reduce disk space usage. [1]
ReferenceIt is an optimization that can reduce storage and processing requirements.
Fragment wantedreduce storage
Judge saidIt correctly states a reduction in storage (disk space) despite omitting processing requirements.
Plain reading"A WITHOUT ROWID table can sometimes reduce disk space usage" is factually true and directly responsive to "what does it reduce"; it is incomplete in omitting processing/performance gains but asserts nothing false.
Skeptical reading"Can sometimes reduce disk space usage" is factually true and directly responsive to "what does a WITHOUT ROWID table reduce?", and it asserts nothing false. The skeptical objection is that it is only half the reference (it omits the processing/performance reduction), but an incomplete-yet-accurate answer to the question asked is not a wrong one, and the hedge "sometimes" matches the SQLite documentation's own wording rather than dodging. Agreeing with the LLM judge, though this is the weakest of the three.
Answer givenAn SQLite index can also reference expressions involving table columns, either directly in the CREATE INDEX statement or by referencing a VIRTUAL generated column [4].
ReferenceAn index can be formed on expressions involving table columns, either by writing an expression in the CREATE INDEX statement or by indexing a generated column.
Fragment wantedformed on expressions
Judge saidThe candidate correctly states that indexes can reference expressions either directly or via a generated column, matching the reference's meaning.
Plain readingIt gives the right answer — an index can be formed on expressions over table columns, written directly in CREATE INDEX or reached via a generated column — and the added word "VIRTUAL" is consistent with SQLite's documented equivalence between an expression index and an index on a virtual generated column, so nothing false is asserted.
Skeptical readingIt names exactly what the question asks for — expressions over table columns, via an expression in CREATE INDEX or via a generated column — and the one addition the reference lacks, the word VIRTUAL, is not a false narrowing: SQLite documents indexing a VIRTUAL generated column as the equivalent of an expression index, and nothing in the answer denies that STORED generated columns can also be indexed. Agrees with the judge.
Answer givenSQLite uses the rowid as the access key for the data in the underlying B-tree storage engine [2]. The rowid serves as the true primary key to look up rows in this storage engine [2].
ReferenceDirectly, via the rowid, which is the fastest lookup available.
Fragment wantedbinary search
Judge saidThe candidate answer conveys that SQLite uses the rowid directly as a key in its storage engine, which matches the reference's description of looking up via rowid directly.
Plain readingThe answer describes the actual mechanism correctly: in a rowid table the rowid is the B-tree key, so the row is reached directly by that key rather than through a secondary index; it omits the "fastest lookup" flourish and the words "binary search" but asserts nothing false.
Skeptical readingI pushed on this one because the answer describes what the rowid IS rather than narrating the lookup, but 'the rowid is the access key for the data in the underlying B-tree' is precisely the how — a direct keyed descent of the table B-tree — and 'the rowid is the true primary key of a rowid table' is verbatim SQLite doctrine, so nothing asserted is false; it merely omits the reference's 'fastest available' and the required 'binary search'. Agrees with the judge.
Answer givenIn SQLite, all reads and writes occur exclusively within transactions. Any command that accesses the database (most SQL commands) will automatically start a transaction if one is not already active, and that transaction is committed when the statement completes [1]. Transactions can also be started manually using BEGIN, and they normally persist until COMMIT or ROLLBACK is issued [1]. Additionally, SQLite supports multiple simultaneous read transactions from separate connections, but only one write transaction at a time [2].
ReferenceNo reads or writes occur except within a transaction. Any command that accesses the database will automatically start one if it is not already in progress.
Fragment wantedno reads or writes
Judge saidCandidate conveys the core relationship that all database access occurs within transactions and auto-start if none is active.
Plain readingIt states both halves of the reference — all reads and writes happen inside a transaction, and any database-accessing command auto-starts one if none is active — and the extra detail (auto-started transactions commit when the statement finishes; many concurrent readers but one writer) is also accurate.
Skeptical readingThe core is right (all access is inside a transaction; any accessing command auto-starts one) and every extra claim I checked also holds — auto-started transactions commit when the statement/last query finishes, manual BEGIN transactions persist until COMMIT or ROLLBACK, and multiple simultaneous read transactions from separate connections with only one write transaction is SQLite's documented concurrency rule; no false assertion rides along. Agrees with the judge.
A living map of modern AI — kept current every morning