Sort a security team's alert bundles into real incidents
Alerts arrive in bundles that look related, and someone must decide which are real and which belong together. This app calls each alert real or false, groups the real ones into incidents and drafts a containment note for the on-call analyst to approve.
PresenterOpens the private repo. Visible to admins only.
For the security operations deskCross-domain · Retail
Why it matters
Today's manual process, and the same job with the app
A security analyst on shift, working a queue of alert bundles for a retailer or any company's IT team.
✕Today's manual process
1Read every alert in the bundle and its clues: addresses, devices, accounts and files.
2Decide which alerts are real, and which a user has already explained away.
3Decide which real ones are one incident, then write the containment note.
4One wrong merge hides a second break-in behind the fix for the first.
Every bundle read and written up manually
✓With the app
1Each alert is read on its own clues, not on how calmly it is worded.
2Each alert is called real or false, so a polite phishing report is still caught.
3Real alerts are grouped into incidents on real links, not a shared address alone.
4One containment note per incident cites a real clue and waits for the on-call analyst to approve it.
The analyst approves every containment step
See it work
One real alert bundle: what the app reads, step by step
Account bcollins is broken into after 18 failed logins, while cbennett logs in from the same address during a conference trip.
Sort a security team's alert bundles into real incidentsReference appBuilt to be shaped to your process
4
1The first alert a login that worked right after a string of failed ones.
2The break-in 18 failed tries on the same account, from the same address.
3A third login a different user, who confirmed a conference trip by phone.
4What does not count one shared address alone does not make these one incident.
For engineers
How it is built, and how we measured it
All fourteen steps of the build are written up, from the business case to running it in your own environment.
Sort a security team's alert bundles into real incidents
A small, forkable project that does one job end to end. Run once for real, and every figure on these pages captured from that run.
PresenterOpens the private repo. Visible to admins only.
The business caseThe problem this solves
A bundle of alerts that look related is two questions, not one, and getting either wrong is expensive in a different direction. Call a genuine phishing report a false positive because it is worded as a calm 'is this legit?' question, and an incident goes unworked. Merge two unrelated alerts into one case because they share a source IP -- the same building's or VPN concentrator's egress address -- and the real second incident is hidden behind the first one's containment action while the analyst's read is spent on a coincidence. A SOC analyst reading a bundle of 2-4 correlated-looking alerts, deciding which are real and which are explained away, deciding which of the real ones are actually the same incident rather than a coincidence, and drafting the containment note -- for every bundle, every shift.
Audience
SOC analysts working an alert queue and deciding what is real and what belongs together, and the people who build triage tooling for them. Every number on these pages came from one real run of this code, not from a vendor page.
The inputThe actual case windows
The corpus is 33 case windows, 0.03 MB (jsonl 1). A real SOC's alert queue names real users, real hosts and real attacker infrastructure -- this repo has never held one and is not able to. There is no public corpus of paired (alert group, adjudicated disposition, confirmed incident grouping) records, for the same reason no public corpus of production payment runs or vendor contracts exists: the interesting cases are exactly the ones nobody can publish. Every 'external' IP in this corpus is drawn from the RFC 5737 documentation ranges (203.0.113.0/24, 198.51.100.0/24), which cannot resolve to anyone's real host; every user, host, domain and hash is invented.
The corpus
The 33 case windowsgenerated from a fixed seed, so no real record, person or institution appears in it.
Where each came fromwritten for this kit rather than collected — the corpus is generated in the kit's own repository, so there is no third-party data in it.
Swap this folder for your own material and the kit is pointed at your case windows. That is the whole change — there is no database to migrate.
One case window, as the model receives itwindows.jsonl · 1 of 33
{"id": "cw001", "window_start": "2026-08-17T06:37:00Z", "on_call_analyst": "Priya Shah", "alerts": [{"alert_id": "ALT-0001", "alert_type": "malware_detection", "description": "Endpoint agent flagged an unsigned binary establishing a reverse shell connection shortly after execution.", "entity": "host:crm-db-05", "indicators": {"file_hash": "d1dc5e74f852a7c2750a5ea08317e789d25ecb0b", "process_name": "sysmaint.exe", "host_ip": "10.20.174"}, "ts": "2026-08-17T06:37:00Z"}, {"alert_id": "ALT-0002", "alert_type": "data_exfil_alert", "description": "Unusual outbound transfer of a large archive to an external IP shortly after the malware alert on the same host.", "entity": "host:crm-db-05", "indicators": {"dest_ip": "203.0.113.106", "bytes_transferred": "1148737880", "source_host": "crm-db-05"}, "ts": "2026-08-17T06:38:00Z"}]}
{"id": "cw002", "window_start": "2026-08-17T07:14:00Z", "on_call_analyst": "Marcus Webb", "alerts": [{"alert_id": "ALT-0003", "alert_type": "malware_detection", "description": "Endpoint agent flagged an unsigned binary establishing a reverse shell connection shortly after execution.", "entity": "host:db-12", "indicators": {"file_hash": "705b10bbfc5641aee6bf795c8840e6f0ab0d4179", "process_name": "sysmaint.exe", "host_ip": "10.20.72"}, "ts": "2026-08-17T07:14:00Z"}, {"alert_id": "ALT-0004", "alert_type": "data_exfil_alert", "description": "Unusual outbound transfer of a large archive to an external IP shortly after the malware alert on the same host.", "entity": "host:db-12", "indicators": {"dest_ip": "203.0.113.221", "bytes_transferred": "3886850958", "source_host": "db-12"}, "ts": "2026-08-17T07:19:00Z"}]}
Abridged — the file continues.
The outcomeWhat a good result looks like
A true_positive/false_positive disposition for every alert in the window, the true positives grouped into the incidents they actually belong to, and one drafted containment recommendation per incident citing a real indicator -- for the named on-call analyst to approve before anything is done.
And when it cannot
Missing a true positive worded like routine correspondence (missed_true_positive), or putting two alerts in one case that gold keeps apart (false_correlation). Measured at 0 of 52 true positives and 0 of 47 gold-different pairs this run, including 0 of 3 and 0 of 8 on the planted trap subsets -- see not_good_enough for why zero on one synthetic run is not the same claim as zero on a real queue.
Where it fitsWhat did work
Every line below is a measured result from this kit's own runs, with the figure that supports it. The headline above is not softened by any of them.
Deciding which alerts in a pre-bundled candidate group are real and which are explained away — the fast tier, over the free keyword floor 100%% disposition accuracy and 0 of 52 missed true positives this run, against the floor's 69.5%% and 4 of 52 -- including 2 of the 3 calmly-worded phishing reports a keyword scan has nothing to catch.
Deciding which true positives are genuinely the SAME incident — the fast tier's case grouping, read pair by pair rather than group by group 0 of 47 gold-different pairs merged this run, including all 8 planted coincidental-indicator pairs, against a floor that merges 8 of 8 by construction. Pair-wise is the honest unit: a partial merge inside a bigger group is still a false correlation for every pair it touches.
What this kit is not — a disposition, a grouping and a drafted recommendation for a named on-call analyst to approve -- never a containment action There is no function in src/triage.py, src/app.py or the UI that locks an account, blocks mail flow or isolates an endpoint. Every recommendation is text.
At a glanceHow the whole thing runs
100%disposition accuracy pct
9,758 msp50, end to end
$5.46per 1,000 case windows · Google Gemini 3 Flash
Run once, for real, on 2026-08-20. Every figure on these pages was captured from that run — nothing is written from intent.
14 steps, grouped by the question that sends you to them rather than by build order. Each tile carries the one figure that step is about, and opens the page behind it.
Sort a security team's alert bundles into real incidents14 steps · 4 questions · run once, for real · 2026-08-20
Every tile is a linkall 14 steps · no step without a page
Should you use this?What you bring, where it stops, and when not to use it
Before you commit an afternoon to this, these are the answers that decide it. Each one is rendered from the record it lives in — and links the page that holds it in full.
What do I have to bring?
Point tools/build_corpus.py's window plan, alert templates and indicator vocabulary at your own alert types, entities and case shapes, or write your own data/windows.jsonl in the same shape -- src/triage.py and src/prompt.py read a window by its own fields and do not care where they came from. The measured 100%% disposition accuracy and 0%% missed_true_positive / false_correlation rates are THIS corpus's alert phrasing (a fixed set of description templates per alert type) and THIS corpus's trap design (3 mundane phishing reports of 9 true-positive phishing alerts; 8 coincidental-indicator windows of the 26 that present a genuine merge-or-don't decision).Corpus lens →
When is this the wrong choice?
Avoid: The free keyword floor for anything but the non-trap windows, where trusting a serious-sounding alert type happens to agree with the correct call -- it is not a competitor, it is the honest floor a model has to clear. That is the case against the best-fitting scenario (“Deciding which alerts in a pre-bundled candidate group are real and which are explained away”). 3 scenarios scored in all, each with its own.Eval lens →
Where does it stop working?
Indicator coverage is assumed complete for every alert. An alert from a security tool this deployment has no feed for arrives with fewer indicators or none, and this kit cannot tell 'nothing to cite' apart from 'the tool never reported it'. 5 recorded failure modes, each from a run rather than a guess.Corpus lens →
What was never verified?
Whether these rates hold on a window meaningfully wider than 4 alerts -- 18 of 33 windows carry 2, 14 carry 3, and exactly ONE carries 4. The pairwise correlation judgement is the part that should degrade first and it was never measured there. 6 items this kit says it could not check.Eval lens →
Can I run this on a model I control?
Yes — any OpenAI-compatible endpoint, including one on your own hardware. The shipped adapter takes its host from BASE_URL and its model from MODEL, so nothing in src/ changes. The published figures come from 1 model on the fast tier, one provider, one key. Prompt lens →
And if it fits — what do I stand up?
3 artifacts with a stated home and a stated egress, and 3 decisions each with what you provision past its ceiling — plus what was not measured. That is the next page, not this one.step 14 — Run it in your environment →
Not asked of this kit — 2 questions: clone (a fresh clone of this kit runs with nothing fetched); judge (nothing here is graded by a model).
Last verified 2026-08-20 — r001-it-sectriage. Every figure on these pages was captured from that run.
Run itHow this reaches your data
Every result on this page was produced by pure code over checked-in files, with no API key — which is why you can read the numbers before anyone spends anything.
Run this on your own data
The pipeline, its eval harness and the runs behind every numberdeployed inside your environment, on your own model endpoints, against your own documents.
The corpus above is the shape, not the limitit is a folder swap, and there is no database to migrate.
Checked before this shipped — A fresh clone with no key configured renders the whole panel -- every window's alerts, indicators and gold grouping, since data/windows.jsonl and data/gold.jsonl are committed, not fetched (python -m src.app). Clicking 'Ask the model' with no API_KEY returns a calm 200 explaining nothing was called. It cannot reproduce a disposition, a score or a dollar figure without a key -- those are what results/eval-r001-it-sectriage.json already committed. The free floor (python -m evals.baseline) and the corpus self-check (python tools/build_corpus.py --verify) both run on a cold clone with no key at all.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
Step 01 of 14Business case
Written for: product manager · Measured during the eval run. Never estimated, never quoted from a vendor page.
In brief
100.0%rows answered
9,758 msp50, end to end
47,346 msp95
2 minclone to first result
What the clock covers. END-TO-END per case window: one HTTP request carrying every alert in the window with its type, entity, timestamp, description and indicators, parsed into a disposition per alert, a case grouping and a drafted recommendation per case. No retrieval step -- loading a window from disk is pure code and costs no network time. Reasoning was left at the provider's default (on), so p50 and p95 both include a reasoning pass on every call; the p95 of 47,346 ms IS one specific window, cw026, the corpus's hardest correlation trap, which spent 22,431 reasoning characters and 5,442 output tokens on it. The spread is real and it is the workload's, not the network's: the fastest call took 1,644 ms.
Current processWhat it replaces
A SOC analyst reading a bundle of 2-4 correlated-looking alerts, deciding which are real and which are explained away, deciding which of the real ones are actually the same incident rather than a coincidence, and drafting the containment note -- for every bundle, every shift.
Where it is not good enough
Four findings, and the first one is the reason to distrust the rest of this page least. FIRST, EVERY HEADLINE NUMBER ON THIS KIT IS 100% OR 0% AND THAT IS A REASON TO VERIFY HARDER, NOT TO TRUST MORE. The corpus is invented (tools/build_corpus.py, seed 20260819) and contains exactly the two failure modes planted on purpose and no others; a real alert stream carries noise and incident shapes this set does not. It is one run of one model -- no repeat, no second tier, no distribution -- and no red-team run was fired. The trap denominators are small and honestly so: 3 mundane-phishing alerts and 8 coincidental-indicator pairs is what a corpus of this scale supports, not a stable rate. What the run DOES establish is separability: the free floor over the identical corpus falls into the correlation trap on 8 of 8 pairs and scores 69.5% disposition accuracy, so a zero here is a real pass rather than a task nothing could fail. SECOND, Reasoning ('thinking') was left at the provider's documented default -- ON -- for this run, and this kit ships NO knob to turn it off: src/adapters/__init__.py takes no thinking parameter at all, unlike rcv-disc's and fin-payrun's. That is worth stating precisely, because it is the OPPOSITE of those kits' finding: there is no discrepancy here between the registered run and the shipped app, since src/app.py reaches the provider down the same adapter with the same defaults. Run and app agree. What they agree ON is expensive: reasoning_chars totals 208,731 across the 33 calls (min 331, max 23,475), which at roughly four characters per token is about 52,183 tokens -- an estimated 95.7% of the run's 54,550 output tokens. A provider bills reasoning as completion tokens, so on the projected rate card that is about 86.9% of this kit's entire dollar cost and the largest single driver of its latency (p50 9758 ms, p95 47346 ms). A forker who wanted that bill cut would have to ADD the adapter parameter rcv-disc already has; nothing in this kit measured what the answers look like with reasoning off. THIRD, MAX_TOKENS is 8192 (src/prompt.py) and it took THREE runs to find that number. At 3000, 7 of 33 calls came back finish_reason='length' and the scorer reported a 34.6% missed-true-positive rate; at 4096, 5 of 33 were still truncated and the rate was still inflated. Both were TRUNCATION ARTEFACTS, not model failures -- at 3000 the disposition accuracy on the alerts that WERE answered was already 100%, and only 61 of 82 alerts got an answer at all. The cause is the reasoning pass: it is billed and bounded as completion tokens, output ran 105-5442 tokens per call against an essentially flat 932-1161 input, and correlating 2-4 alerts against each other is heavier reasoning than a per-record judgement. The published run's worst call used 5,442 of 8,192, so the shipped ceiling has real headroom rather than sitting just past the observed maximum. FOURTH, two structural limits stated in the corpus's own documentation rather than found by this run: indicator coverage is assumed complete for every alert, so an alert from a security tool this deployment has no feed for -- arriving with fewer indicators or none -- is indistinguishable to this kit from an alert whose evidence is merely thin; and the case-window boundary itself is authored, not computed, so which alerts get bundled together in the first place is an operator-tuned entity/time correlation window upstream of everything measured here.
One question, end to end. Colour on the fan-in is the corpus format; a red hook marks every station with a recorded failure, and the words — every sub line, failure and report field — are in the ledger beneath. Every station is a link — click one to jump to its row, and each row links to the lens that explains it.
The free keyword-and-shared-indicator baseline (evals/baseline.py, which groups on any shared entity or indicator VALUE and then spreads the disposition by association) scores 69.5 pct disposition accuracy and merges all 8 planted coincidental-indicator pairs -- a 100 pct false-correlation rate on exactly the pairs this kit's model kept apart.
⚠︎ THE PUBLISHED RUN IS THE THIRD ATTEMPT: at a 3000-token reply ceiling 7 of 33 calls returned finish_reason=length and the scorer reported 34.6 pct missed true positives off only 61 of 82 alerts answered, with 100 pct accuracy on the ones that were. That was truncation, not judgement. No red-team run exists for this kit -- this footer names that absence rather than a resistance rate it does not have.
The swap seams
Seam
File
What changes
the model
src/adapters/__init__.py
.env plus the same run again. Any OpenAI-compatible provider or Anthropic.
the reply ceiling
src/prompt.py
MAX_TOKENS, and on this workload it is not a formality -- 3000 and 4096 both truncated and both produced a fake missed-true-positive rate. Changing it is a paid re-run, not a re-score; see Business.not_good_enough.
whether reasoning is enabled
src/adapters/__init__.py
NOTHING TODAY, AND THAT IS THE FINDING. This kit ships no thinking parameter, so reasoning is whatever the provider defaults to on both the run and the live app. Making it a seam means adding the kwarg rcv-disc's adapter already carries -- an estimated 86.9% of this kit's projected cost sits behind a knob that does not exist yet.
the corpus
tools/build_corpus.py
Point its window plan, alert templates and indicator vocabulary at your own alert stream, or write your own data/windows.jsonl in the same shape -- src/triage.py and src/prompt.py read a window by its own fields (alerts[].alert_id/alert_type/entity/ts/description/indicators) and do not care where they came from.
the trap fractions
tools/build_corpus.py
Set by the five-pattern window plan, not one global -- a new corpus states its own pattern counts in the same shape, and evals/scoring.py takes the trap alert ids and trap pairs as explicit arguments read from gold rather than re-guessing which windows look like traps.
Components
Component
File
Role
corpus generator
tools/build_corpus.py
Generates 33 case windows and 82 alerts from a fixed seed (SEED=20260819) across five alert types, in five window patterns (7 correlated_tp, 6 uncorrelated_fp, 7 mixed_tp_fp, 5 multi_incident, 8 false_correlation_trap). Gold is DERIVED, never authored beside the data: every alert is tagged at construction time with a disposition and a case_id, and derive_gold() turns those per-alert facts into the alert_dispositions dict and the case_groups list -- used both to write data/gold.jsonl and, independently, by --verify re-reading both files from disk.
the prompt
src/prompt.py
The three-part instruction (disposition, correlation, recommendation), the output schema and MAX_TOKENS, declared once and read from here by build() and parse(). States both failure modes to the model as ordinary SOC judgement -- a calm report is not evidence of nothing, a shared address is not evidence of one incident -- rather than as hints about this corpus.
the AI layer
src/triage.py
Loads one case window, calls the model ONCE with every alert in it, and parses the reply into dispositions, case groups and recommendations. One call per WINDOW, never per alert -- that is what makes correlation possible at all, and it is the whole difference from ops-triage's per-window judging, which never compares one alert against another. Never executes containment: there is no function here, in src/app.py or in the UI that locks an account, blocks mail flow or isolates an endpoint.
the model adapter
src/adapters/__init__.py
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Records finish_reason and reasoning_chars beside the text, which is the only reason the MAX_TOKENS truncation finding was diagnosable without spending again. Takes NO thinking parameter -- see Business.not_good_enough.
spend control
src/budget.py
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose.
the app
src/app.py
Static files and three JSON endpoints (/api/state, /api/window, /api/triage) on the standard library, port 8792. Renders every window, its alerts and the gold answer with no key configured; only /api/triage calls a provider, and with no API_KEY it returns a calm 200 saying so. Reaches the provider down the same adapter with the same defaults the registered run used -- no reasoning-setting discrepancy, unlike rcv-disc and fin-payrun.
the scorer
evals/scoring.py
Four things scored, kept apart, never folded into one accuracy number: disposition accuracy, missed_true_positive (the expensive direction for classification), false_correlation (the expensive direction for grouping) and citation_validity. false_correlation is counted over PAIRS rather than groups, because a partial merge inside a bigger group is still a false correlation for every pair it touches and no group-vs-group comparison describes it.
free baseline
evals/baseline.py
A keyword-and-shared-indicator rule written the way a person free-texting a quick triager would write one: trust the three detector types that sound serious, keyword-match the two a human reports, then group anything sharing an entity or ANY raw indicator value and spread the disposition by association. Not weakened for the demonstration -- it is what 'same IP, must be one incident' looks like as code, and it is correct on every window carrying no trap.
Where it breaks at scale
Not on window count -- each window is independent and one call per window is linear. It breaks three other ways. FIRST, ON WINDOW WIDTH, WHICH IS THE ONE THIS CORPUS CANNOT SPEAK TO AT ALL. Every window here is 2 to 4 alerts, and the distribution is narrower than even that suggests: 18 windows carry 2 alerts, 14 carry 3, and exactly ONE carries 4. A window with meaningfully more alerts than this corpus's 4-alert maximum was never measured, and the pairwise correlation judgement is the part that should be expected to degrade first, since the number of merge-or-don't decisions grows with the square of the alert count while the reasoning budget does not. SECOND, ON THE REASONING CEILING, and here the kit has already been bitten once: output tokens ran 105 to 5,442 per call against a flat 932-1,161 input, and MAX_TOKENS had to be raised twice before it stopped truncating. 8,192 clears this corpus's worst call by 2,750 tokens, but a bigger window would spend more reasoning, not less. THIRD, ON THE STRUCTURAL ASSUMPTION: the case-window boundary is authored upstream, so a real deployment's own entity/time correlation window governs what a model is ever given the chance to get right -- too loose over-merges before any model sees the alerts, too tight splits one incident across two windows that are never compared.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
Step 03 of 14App UI
Written for: anyone, in 5 seconds · Screenshots of the running app taken during the run.
SuccessesWhen it works
The panel before any call: 33 case windows, 82 alerts, and the two planted traps counted on their own tiles (3 calmly-worded true-positive phishing reports, 8 coincidental-indicator merge traps). Window cw001 is open, showing both alerts with their full indicator sets and the gold case grouping -- the page states what 'correct' means before anyone spends a call.successOpen full size →
LimitsWhen it does not
A report showing only wins is an advert. This one is required, and the validator now refuses a kit that omits it.
Window cw026 -- the hardest correlation trap in the corpus -- after pressing 'Ask the model' with no API_KEY configured: a calm 200 explaining nothing was called, rather than an error. The trap is visible in the same frame: all three alerts share source_ip=198.51.100.177, two of them share device_id=dev-c2277, and gold still keeps ALT-0063 out of the case. This is the honest failure state, not a staged one.failureOpen full size →
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
Step 04 of 14Corpus
Written for: "can I point this at mine?" · Measured off the corpus the run actually used.
In brief
33case windows
0.03 MiBjsonl 1
p50 656chars per characters per window's assembled alert block
$0.00setup · 0.0s
How it is cutWhat one characters per window's assembled alert block is
One case window, every alert in it included whole -- no train/test split, every window judged once, in one call. 18 windows carry 2 alerts, 14 carry 3, 1 carries 4.
SetupWhat the setup figure measured
No index is built. src/app.py finds a window by scanning T.windows() for a matching id -- 33 rows, a linear scan, not a search structure -- and src/triage.py sends that window's alerts whole into the one call. build_seconds and build_cost_usd are both zero because there is no index-build step, not because one ran for free.
LicenceLicence
MIT (this repository's own licence)
Bring your ownBring your own case windows
Point tools/build_corpus.py's window plan, alert templates and indicator vocabulary at your own alert types, entities and case shapes, or write your own data/windows.jsonl in the same shape -- src/triage.py and src/prompt.py read a window by its own fields and do not care where they came from. The two-value disposition vocabulary in src/prompt.py assumes exactly true_positive and false_positive; a third value (a 'needs enrichment' state, say) is a prompt.py AND tools/build_corpus.py change, not a data-only change.
⚠︎ And what stops being true when you do: The measured 100%% disposition accuracy and 0%% missed_true_positive / false_correlation rates are THIS corpus's alert phrasing (a fixed set of description templates per alert type) and THIS corpus's trap design (3 mundane phishing reports of 9 true-positive phishing alerts; 8 coincidental-indicator windows of the 26 that present a genuine merge-or-don't decision). A real SOC's own alert text, and a real network's own reasons for two users to share an egress address, are both unmeasured by this run.
What breaks it
Indicator coverage is assumed complete for every alert. An alert from a security tool this deployment has no feed for arrives with fewer indicators or none, and this kit cannot tell 'nothing to cite' apart from 'the tool never reported it'.
The case-window boundary is authored, not computed. Which alerts get bundled into one 2-4 alert candidate group is an operator-tuned entity/time correlation window upstream of this kit, and nothing here measures that choice.
Every window carries at most one of the two named traps, to keep each failure mode legible on its own. A real SOC queue can carry both kinds of trouble in the same bundle at once.
The false-correlation trap's coincidental shared indicator is ALWAYS a source IP in this corpus. A real deployment sees the same failure shape from a shared device fingerprint, a shared network segment or a shared vendor mail relay; this corpus tests the mechanism with one concrete instrument of it.
The trap denominators are small -- 3 mundane-phishing alerts and 8 coincidental-indicator pairs. That is what a corpus of this scale supports honestly, not a stable rate.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
Step 05 of 14Prompt
Written for: engineer · Logged from the run, verbatim.
In brief
The assembled prompt in full, and its real decomposition. Every part below occurs in the prompt that was actually sent, and the token counts sum to the recorded input total — but they are the input total apportioned by each part’s share of the characters, not a per-part reading from the provider’s counter. The total is measured; the split is arithmetic, and a part carrying denser text really costs more of it than this says.
AssemblyHow the prompt is assembled
Part
Characters
Tokens
system
2,778
805
alerts
560
162
Total
967
This is the cost lesson as arithmetic: of the 967 tokens assembled, 805 are instructions — 83% of the prompt, the largest single part. A per-request bill follows whichever part is biggest, and on this kit that is the one named here.
Logged verbatim by importing src.prompt and src.triage directly and calling prompt.build(window) for cw001 (not retyped) -- the literal system and user content src/triage.py::check() sends. Per-part token counts are a proportional estimate over character share (system 2778 / alerts 560 chars, 3338 total) applied to this call's own total of 967 input tokens -- the provider reports only the call's total, matching lenses.LLM.tokens.input exactly, never a per-segment split. The 'alerts' part covers only the alert block src/prompt.py::_alert_block() builds; the surrounding 'Case window ... Return the JSON object ...' instruction text is part of the user message actually sent (see prompt_verbatim) but is not a separately named segment in the kit's own code.
Full promptThe prompt in full
The verbatim prompt, as sent
the prompt, as sent
You are a SOC triage analyst. You are shown one case window: a small group of 2-4 security alerts that were bundled together because they occurred close in time and/or share an entity. Your job has three parts, in one reply:
1. DISPOSITION -- for every alert in the window, decide true_positive (a real security event) or false_positive (benign, correctly explained away). Judge each alert on its own evidence.
- A calm or routine-sounding report is not evidence of a false positive. Genuine phishing reports are often written as an ordinary 'is this legit?' question, not an alarm -- judge the sender domain, the link and the ask, not the tone.
- A loud description is not automatically a true positive either -- read whether it is actually explained (a confirmed backup window, a confirmed travelling user, an approved software request).
2. CORRELATION -- among the alerts you called true_positive, decide which ones are genuinely the SAME incident and which are separate. Two alerts belong in the same case only when the evidence connects them causally, not just superficially.
- Do not merge two alerts into one incident merely because they share an entity, a close timestamp, or one matching indicator value. A shared source IP can mean one attacker, or it can mean two unrelated users behind the same building's network or the same VPN concentrator -- read what the shared value actually is before treating it as a link.
- A true positive with no genuine partner is its own case, alone.
- Never place a false_positive alert into any case -- a case exists only where there is a real incident.
3. RECOMMENDATION -- for every case with at least one true-positive alert, draft one short containment recommendation (account lockout, mail-flow block, or endpoint isolation, whichever fits). Cite the specific indicator(s) that justify it, in the exact form they appear in the alert data (name=value). Never cite an indicator that is not actually present on an alert in that case, and never invent one.
YOU NEVER EXECUTE CONTAINMENT. You only draft the recommendation. The named on-call analyst shown to you reviews and approves it before anything is done -- your recommendation text should read as a draft awaiting that approval, not as an action taken.
Reply with ONLY a JSON object, no other text, with exactly these three keys:
"alert_dispositions": {"<alert_id>": "true_positive"|"false_positive", ... one entry per alert in the window}
"case_groups": [[<alert_id>, ...], ...] -- one list per incident, true-positive alert ids only, every true positive in exactly one group
"recommendations": [{"case": [<alert_id>, ...], "action": "<one or two sentences>", "citations": ["<indicator_name>=<value>", ...]}, ...] -- one entry per case_groups group
Case window cw001, starting 2026-08-17T06:37:00Z. On-call analyst: Priya Shah.
Alerts in this window (ALT-0001, ALT-0002):
ALT-0001 [malware_detection] entity=host:crm-db-05 06:37:00
Endpoint agent flagged an unsigned binary establishing a reverse shell connection shortly after execution.
indicators: file_hash=d1dc5e74f852a7c2750a5ea08317e789d25ecb0b, process_name=sysmaint.exe, host_ip=10.20.174
ALT-0002 [data_exfil_alert] entity=host:crm-db-05 06:38:00
Unusual outbound transfer of a large archive to an external IP shortly after the malware alert on the same host.
indicators: dest_ip=203.0.113.106, bytes_transferred=1148737880, source_host=crm-db-05
Return the JSON object described in the system message. Every one of these alert ids must appear in alert_dispositions, exactly once: ALT-0001, ALT-0002
Raw responseThe raw response
The raw response, before any parsing
the response, unparsed
{
"alert_dispositions": {
"ALT-0001": "true_positive",
"ALT-0002": "true_positive"
},
"case_groups": [
[
"ALT-0001",
"ALT-0002"
]
],
"recommendations": [
{
"case": [
"ALT-0001",
"ALT-0002"
],
"action": "Isolate host crm-db-05 immediately and block outbound traffic to the external destination to contain the reverse shell and stop further data exfiltration. Preserve the endpoint for forensic collection.",
"citations": [
"file_hash=d1dc5e74f852a7c2750a5ea08317e789d25ecb0b",
"process_name=sysmaint.exe",
"dest_ip=203.0.113.106",
"source_host=crm-db-05"
]
}
]
}
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
Step 06 of 14Evals
Written for: the skeptic · The harness output, with as_of / model version / dataset version / verified_by.
In brief
Sort a security team's alert bundles into real incidents — 33 case windows. One model answered, and every answer was then graded one way — by == against a fixed gold value, which is a measurement too and only as good as the gold behind it.
The rulerHow it was graded
The grader is evals/scoring.py and it is pure code: exact match on each alert's disposition, a pair-wise comparison for correlation, and an exact (indicator_name, value) lookup for every citation. No LLM judge anywhere in the path, and no retrieval step to score.
33case windows
33source documents
1model tier
1grading method
MeasurementsWhat was measured
COUNTED82 / 82disposition accuracy pct — alerts with a parsed dispositionDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED82 / 82answered pct — alerts askedDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 52missed true positive rate pct — gold true-positive alertsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 3missed true positive trap rate pct — planted mundane-phishing alertsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 47false correlation rate pct — gold-different alert pairsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED0 / 8false correlation trap rate pct — planted coincidental-indicator pairsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
COUNTED109 / 109citation validity pct — citations drafted across 35 recommendationsDecided by == against a fixed gold value. No model graded this and it reproduces to the digit — which makes it checkable, not necessarily right: it is only as good as the gold behind it.
NOT YET KNOWN—A person confirmed the gold is rightThe grader is == and needs no confirming. The gold does: it is machine-derived from the registry's structured modules, and every rate on this page rests on that derivation being faithful to what the prose actually says. Nobody has read a document against its gold record by hand.
Two words carry this page: Counted is deterministic and nobody's opinion and Not yet known is printed blank rather than filled with something plausible. There is no Judged row here and that is the point: nothing on this page is a model's opinion about another model. The row most worth having is currently the empty one.
How the method was validated
The scorer is pure code, exact match against a mechanically-derived gold set: every alert is tagged with a disposition and a case_id at generation time and derive_gold() computes the aggregate dispositions and case groups from those facts, with tools/build_corpus.py --verify re-deriving every row independently against the files actually on disk and asserting zero drift, no false positive inside any case group, and every planted trap_pair genuinely gold-separate.
Grading costWhat it costs
Every dollar on these pages is a MEASURED token count multiplied by a named vendor's PUBLISHED rate -- a projection onto a rate card, not a bill anyone paid.
Priced at
Per 1M in / out
One case window
1,000 case windows
Share that is the prompt
Google Gemini 3 Flash the cheapest model Google publishes a rate for, and the card every kit in this series projects onto so the rows compare across use cases
$0.50 / $3.00
$0.005457
$5.46
9%
Same work, 1× the bill
The same case windows, the same tokens — only the rate card changed. And on that card about 9% of what you pay is the prompt this pipeline sends, not the answer it writes.
Whether the provider's reasoning pass runs at all -- and today there is no lever, which is the finding. src/adapters/__init__.py takes no thinking parameter, so both the registered run and the live app get whatever the provider defaults to. Adding the kwarg rcv-disc's adapter already carries would put an estimated 86.9% of this kit's cost under control; nothing here has priced the two configurations against each other, because only one of them can currently be run.
Rates checked 2026-08-18. The provider that actually ran r001 publishes no rate card this repo commits, so nothing here is what was actually paid -- the real spend for this kit's build is recorded in the commit history and the shared call ledger, not on this page. Reasoning was left at provider default (on) with no way to disable it, so the projected figure prices a reasoning pass that a reader on another provider may or may not be billed for.
Grading unitWhat the grading figure prices
freegrading cost, as measured
Grading calls nothing. Disposition accuracy, missed_true_positive, false_correlation and citation_validity are all pure code over the committed run file, so a forker re-scores this run -- or the free floor -- for $0.00 and needs no key.
The gradersOne way to grade, and why it is the only one
The floor is correct on every window that carries no trap, and it fails exactly where the corpus means it to. It over-merges 8 of 8 planted coincidental-indicator pairs -- 100%% -- because a rule that groups first by raw similarity and only then asks whether the group looks bad will drag an unrelated alert into a real incident's case the moment it shares one indicator value. It misses 2 of 3 mundane-phishing trap alerts (the third is rescued by a loud campaign neighbour it is genuinely correlated with, in cw006). Its 19 of 47 false correlations and 30.5 points of lost disposition accuracy are not a strawman weakening: trusting malware/exfil/brute-force alert TYPES outright is what a rotation under time pressure does, and 'same IP, must be one incident' is what that looks like as code. It also misses a quiet, no-precursor account-takeover login in cw020 and cw024 -- a real gap in the floor that NEITHER named trap set out to test, printed by the baseline rather than smoothed over.
Grader
Cost / 1k
Data leaves
Same answer twice
Result
Disposition exact match, pair-wise false correlation, and citation-existence checks Does each alert's true_positive/false_positive call match gold? Does the model avoid calling a gold true positive a false positive (missed_true_positive), overall and on the 3 deliberately mundane-worded phishing reports? Does it avoid putting two alerts in one case that gold keeps apart (false_correlation), overall and on the 8 planted coincidental-indicator pairs? Does every citation name an indicator that actually exists on an alert in that case?
$0.00
no
yes
the fast tier 100.0% disposition accuracy · 2 more measured on each run
The row opens its own page: how the test was run, one real row, the formula with every iteration of it, and the analysis.
Why this set cannot separate them
Yes, and on both axes this corpus was built to test. The free floor and the fast tier are scored by the identical function over the identical 33 windows, and they separate cleanly: 69.5%% vs 100%% disposition accuracy, 4 vs 0 missed true positives, 19 vs 0 false correlations, and on the planted subsets 2 of 3 vs 0 of 3 and 8 of 8 vs 0 of 8. A corpus where a zero could not be distinguished from a task nothing could fail is a corpus that proves nothing; this one convicts the floor on exactly the windows it was designed to convict it on.
Choosing oneWhen to reach for the model at all
The comparison is only useful if it ends in a choice, and on this kit the choice is not which grader — there is one, and it is ==. It is whether to call a model for this field at all, and the answer is conditional.
If your situation is
Use
Because
And avoid
Deciding which alerts in a pre-bundled candidate group are real and which are explained away
the fast tier, over the free keyword floor
100%% disposition accuracy and 0 of 52 missed true positives this run, against the floor's 69.5%% and 4 of 52 -- including 2 of the 3 calmly-worded phishing reports a keyword scan has nothing to catch.
the free keyword floor for anything but the non-trap windows, where trusting a serious-sounding alert type happens to agree with the correct call -- it is not a competitor, it is the honest floor a model has to clear.
Deciding which true positives are genuinely the SAME incident
the fast tier's case grouping, read pair by pair rather than group by group
0 of 47 gold-different pairs merged this run, including all 8 planted coincidental-indicator pairs, against a floor that merges 8 of 8 by construction. Pair-wise is the honest unit: a partial merge inside a bigger group is still a false correlation for every pair it touches.
trusting the grouping on a window meaningfully wider than this corpus's 4-alert maximum -- 18 of 33 windows carry only 2 alerts and exactly one carries 4, so the pairwise judgement was never measured at scale. See Architecture.breaks_at_scale.
What this kit is not
a disposition, a grouping and a drafted recommendation for a named on-call analyst to approve -- never a containment action
There is no function in src/triage.py, src/app.py or the UI that locks an account, blocks mail flow or isolates an endpoint. Every recommendation is text.
wiring this straight into a containment API. Also avoid reading it as a correlation ENGINE: it judges the bundle it is handed and never decides which alerts get bundled in the first place, which is the upstream choice that governs what it can get right at all.
This pipeline was also attacked, which is a different question from whether its answers are right: see Threat model for the injection experiment, its two gates and what it does not prove.
Failure causesHow it fails
Cause
What it means
Rows
A real example
correlated_tp
One genuine incident; every alert in the window is a true positive and they belong together
7
cw001: a malware_detection on host crm-db-05 (unsigned binary opening a reverse shell) followed a minute later by a data_exfil_alert on the same host (a large archive to an external IP). Gold and model both: both true_positive, both in one case. The model…
uncorrelated_fp
Two or three false positives with nothing in common -- the correct answer is no case at all
6
cw008: alerts a keyword floor trusts on type alone. Gold: every alert false_positive, case_groups empty. The model returned no case; the floor called ALT-0017 a true positive because its alert type sounds serious -- one of its 21 type-trusted false positives.
mixed_tp_fp
One real, solo incident buried among one or two unrelated false positives
7
cw015: ALT-0032 is a genuine phishing report worded as 'just double-checking' an invoice email with slightly different payment details -- one of the 3 planted mundane-phishing trap alerts. Gold: true_positive, its own case. The keyword floor called it…
multi_incident
Two genuinely separate real incidents in one window, sharing no entity or indicator
5
cw021-cw025: each window holds two real incidents that must stay apart. Gold keeps them in separate case_groups and the model did too -- these are the windows where the failure would be UNDER-merging as much as over-merging, and neither occurred.
false_correlation_trap
A real incident and an unrelated alert sharing one coincidental indicator -- must NOT become one case
8
cw026: a brute_force and its successful login (user bcollins) plus a benign travelling-user login (user cbennett) -- all three sharing source_ip=198.51.100.177. Gold: [ALT-0061, ALT-0062] one case, ALT-0063 a false positive in no case. The floor merged all…
What we could NOT verify
Whether these rates hold on a window meaningfully wider than 4 alerts -- 18 of 33 windows carry 2, 14 carry 3, and exactly ONE carries 4. The pairwise correlation judgement is the part that should degrade first and it was never measured there.
Whether the same figures repeat. One run, one model, no repeat -- there is no band, only a point.
Whether reasoning being ON changed the answers. It was left at the provider default on every call and this kit ships no way to disable it, so the reasoning-off configuration has never been run.
Whether the two traps interact. Every window carries at most one of them by construction; a real queue can carry both in the same bundle.
Whether the prompt-level boundaries hold against hostile alert text. No red-team run exists for this kit.
Whether an alert with missing or absent indicators is handled sensibly -- every alert in this corpus carries 2-4 indicators by construction, and the kit cannot tell 'nothing to cite' from 'the tool never reported it'.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
Step 07 of 14Unit cost
Written for: whoever signs off on the bill · Measured over the eval run, per model. Never quoted from a vendor page.
In brief
Tier
Input tokens
Output tokens
p50 latency
Google Gemini 3 Flash
the fast tier
996.73
1,653.03
9,758 ms
$0.005457
The token counts are measured on a real run and belong to this pipeline — which sections are selected, the field schema, the prompt — so they hold wherever you run it. The dollars are those counts multiplied by each vendor's published rate, checked 2026-08-18. Nobody paid these particular bills; they are what the same work would cost you on a card you can sign up to today.
Identical work in both rows — the same query, the same tokens. Only the price list changed. Which is why the token count is the measurement and the dollar is arithmetic you can redo on your own card.
Answering vs gradingAnswering and grading are two separate bills
The keyword-and-shared-indicator baseline (evals/baseline.py) and the scorer (evals/scoring.py) are both pure code and cost $0.00 to run against either result set. The figure above is this run's own 32892 input / 54550 output tokens priced at Google Gemini 3 Flash's published rate, the same basis cost_per_query_usd uses, not a second, larger spend.
Cost driversWhat actually moves the bill
The provider's reasoning pass, and it is not a rounding error -- it is the bill. reasoning_chars totals 208,731 across the 33 calls, roughly 52,183 tokens at four characters each, an estimated 95.7% of the 54,550 output tokens. Output tokens are 90.9% of the projected cost, so an estimated 86.9% of this kit's entire dollar figure is reasoning nobody asked for and this kit ships no knob to decline.
Every alert in the window with its description and 2-4 indicators, plus the fixed system prompt -- essentially flat per call (input tokens ranged 932-1161 across all 33 calls), the floor every call pays regardless of how many alerts the window holds.
The fixed system prompt (2778 characters -- the three-part instruction, the two pieces of general triage guidance and the output schema) is sent in full on every call: 83.2% of the system+alerts character total on the published call.
Your volumeWhat it costs at your volume
Linear in windows: each call is independent and self-contained, with no shared context or retrieval step to amortise. This run's 33 windows cost about $0.1801 projected onto Google Gemini 3 Flash's published rate, so ten times the set is about $1.801 on the same rate and the same reasoning-on configuration -- arithmetic on the measured per-call rate, not a second run. It is linear in WINDOWS and not in ALERTS: a wider window costs more per call, and how much more was never measured past four alerts.
Where pricing changes shape
Your return, with your numbers
Volumealert bundles triaged per shift/day -- this run triaged 33 windows (82 alerts) in one pass
What it replacesa SOC analyst reading a bundle of correlated-looking alerts, calling each one real or explained, deciding which of the real ones are the same incident, and drafting the containment note
Time saved per itemnot measured here -- depends on how long a manual bundle review takes at the reader's own SOC
We publish the inputs, not a return: a return depends on your labour cost and your volume. The ROI dashboard takes it from here, and AI Costing explains the math.
Model choiceWhy this model
The fast tier was the only model run against this corpus. No second tier was run -- unlike data-reconcile's fast/reasoning comparison -- so this page prices one model, not a trade-off. Note the awkward shape that produces here: a 'fast tier' whose p95 is 47 seconds and whose output-token average is larger than its input, because the provider's default reasoning pass is doing most of the work. See Eval.could_not_verify for what a second run would need to answer.
Other modelsOn other models
This is arithmetic, not a run. No model below was executed against the labelled set, and no accuracy is claimed for any of them. What transfers is the input volume: retrieval is model-independent, so every model would receive the same passages.
32,892input tokens · this run
54,550output tokens
—not priced — no committed card for the provider that ran it
The exact work behind every number on these pages: 33 case windows triaged, 82 alerts dispositioned, 47 gold-different pairs and 109 citations scored by pure code. This run (r001-it-sectriage) answered 82 of 82 dispositions -- the row this table prices, the same one Cost.cost_by_model[0] uses.
Model
Provider
One eval pass
This whole run
Per 1,000 queries
Rates as of
gpt-5-6-luna
OpenAI
$0.072
$0.072
$2.18
2026-09-12
gemini-3-flash
Google
$0.180
$0.180
$5.46
2026-09-18
gemini-3-8-flash
Google
$0.229
$0.229
$6.95
2026-09-18
llama-5
Meta
$0.273
$0.273
$8.27
2026-09-18
claude-haiku-4-5
Anthropic
$0.306
$0.306
$9.26
2026-09-12
grok-4-5
xAI
$0.393
$0.393
$11.91
2026-09-18
grok-4-6
xAI
$0.393
$0.393
$11.91
2026-09-18
claude-sonnet-5
Anthropic
$0.611
$0.611
$18.52
2026-09-12
gemini-3-1-pro
Google
$0.720
$0.720
$21.83
2026-09-18
gpt-5-6-terra
OpenAI
$0.720
$0.720
$21.83
2026-09-12
gpt-5-6-sol
OpenAI
$1.223
$1.223
$37.05
2026-09-12
claude-opus-4-8
Anthropic
$1.528
$1.528
$46.31
2026-09-12
claude-opus-5
Anthropic
$1.528
$1.528
$46.31
2026-09-12
claude-fable-5
Anthropic
$3.056
$3.056
$92.62
2026-09-18
claude-fable-5-1
Anthropic
$3.056
$3.056
$92.62
2026-09-18
gpt-6-astra
OpenAI
$3.056
$3.056
$92.62
2026-09-17
Read this against the numbers above
REASONING WAS NOT, AND CANNOT CURRENTLY BE, DISABLED FOR THIS WORKLOAD. Every row below prices THIS run's own token counts, an estimated 95.7% of whose output is a provider-side reasoning pass. This kit's adapter takes no thinking parameter, so a forker's real bill depends on whether their chosen model reasons by default and whether they add the knob.
NO QUALITY IS IMPLIED. Only the model that produced Eval.scores has been scored against this corpus -- every row here is a price, not a recommendation. A cheaper model that truncates its reply on this workload would score badly for a reason that has nothing to do with its judgement, which this kit has already demonstrated twice against itself.
Rates are read from build/facts/models.json, each row carrying its own as_of and source. A price is the fastest-ageing fact on this page.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
The codeEvery file, and what it is for
Eight modules, no framework, and almost no dependencies. The pipeline is stdlib Python, because a kit is meant to be read and changed by someone who cloned it ten minutes ago — every dependency is a thing they have to understand before they can touch the part they came for.
Five of these are swap seams: the points you edit to make the kit yours. Each is marked, and each block below is the module’s real shape, read out of the file rather than written here.
The files on this page are not in the public GitHub catalog.
tools/build_corpus.pycorpus generator — a swap seam
Generates 33 case windows and 82 alerts from a fixed seed (SEED=20260819) across five alert types, in five window patterns (7 correlated_tp, 6 uncorrelated_fp, 7 mixed_tp_fp, 5 multi_incident, 8 false_correlation_trap). Gold is DERIVED, never authored beside the data: every alert is tagged at construction time with a disposition and a case_id, and derive_gold() turns those per-alert facts into the alert_dispositions dict and the case_groups list -- used both to write data/gold.jsonl and, independently, by --verify re-reading both files from disk.
You change it to: Set by the five-pattern window plan, not one global -- a new corpus states its own pattern counts in the same shape, and evals/scoring.py takes the trap alert ids and trap pairs as explicit arguments read from gold rather than re-guessing which windows look like traps.
tools/build_corpus.py
# Generate the case windows and gold labels this kit checks against, from a fixed seed.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DATA = os.path.join(HERE, "data")
SEED = 20260819
START = datetime(2026, 8, 17, 6, 0, 0, tzinfo=timezone.utc)
ANALYSTS = ["Priya Shah", "Marcus Webb", "Dana Okafor", "Tomas Rivera"]
USERS = ["jrivera", "asingh", "mchen", "dokafor", "lferreira", "kwalsh", "tnguyen", "rpatel",
HOSTS = ["web-03", "db-12", "fin-app-02", "vpn-gw-01", "mail-relay-04", "build-agent-07",
ALERT_TYPES = ("phishing_report", "suspicious_login", "malware_detection", "data_exfil_alert",
EXT_IP_A, EXT_IP_B = "203.0.113", "198.51.100"
src/prompt.pythe prompt — a swap seam
The three-part instruction (disposition, correlation, recommendation), the output schema and MAX_TOKENS, declared once and read from here by build() and parse(). States both failure modes to the model as ordinary SOC judgement -- a calm report is not evidence of nothing, a shared address is not evidence of one incident -- rather than as hints about this corpus.
You change it to: MAX_TOKENS, and on this workload it is not a formality -- 3000 and 4096 both truncated and both produced a fake missed-true-positive rate. Changing it is a paid re-run, not a re-score; see Business.not_good_enough.
src/prompt.py
# Assemble the one prompt this kit sends, and parse the one reply it gets back.
DISPOSITIONS = ("true_positive", "false_positive")
MAX_TOKENS = 8192
SYSTEM = (
DEFAULT_PROMPT = "v1"
SYSTEMS = {"v1": SYSTEM}
def _alert_block(a):
def build(window, prompt=DEFAULT_PROMPT):
def _strip_fence(text):
def parse(raw):
src/triage.pythe AI layer
Loads one case window, calls the model ONCE with every alert in it, and parses the reply into dispositions, case groups and recommendations. One call per WINDOW, never per alert -- that is what makes correlation possible at all, and it is the whole difference from ops-triage's per-window judging, which never compares one alert against another. Never executes containment: there is no function here, in src/app.py or in the UI that locks an account, blocks mail flow or isolates an endpoint.
src/triage.py
# Triage one case window: classify, correlate, draft -- one model call, done.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
WINDOWS = os.path.join(HERE, "data", "windows.jsonl")
GOLD = os.path.join(HERE, "data", "gold.jsonl")
MAX_TOKENS = P.MAX_TOKENS
def windows():
def load_gold():
def check(cfg, window, complete=None, prompt=P.DEFAULT_PROMPT):
src/adapters/__init__.pythe model adapter — a swap seam
SEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Records finish_reason and reasoning_chars beside the text, which is the only reason the MAX_TOKENS truncation finding was diagnosable without spending again. Takes NO thinking parameter -- see Business.not_good_enough.
You change it to: NOTHING TODAY, AND THAT IS THE FINDING. This kit ships no thinking parameter, so reasoning is whatever the provider defaults to on both the run and the live app. Making it a seam means adding the kwarg rcv-disc's adapter already carries -- an estimated 86.9% of this kit's projected cost sits behind a knob that does not exist yet.
src/adapters/__init__.py
# SEAM 1 — the model. Swapping provider or model is .env plus the same run again.
class AdapterError(RuntimeError):
TRANSIENT = {408, 429, 500, 502, 503, 504}
RETRIES = 4
def _post(url, headers, payload, timeout=120):
def openai_compatible(cfg, system, user, max_tokens):
def anthropic(cfg, system, user, max_tokens):
PROVIDERS = {"openai-compatible": openai_compatible, "anthropic": anthropic}
def complete(cfg, system, user, max_tokens=1024):
src/budget.pyspend control
An append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose.
src/budget.py
# A cap on live calls, shared by every kit on this machine. No dependency, no service.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
KIT = os.path.basename(HERE)
ROOT = os.path.dirname(os.path.dirname(HERE))
SHARED_ENV = os.path.join(ROOT, ".env")
LEDGER = os.path.join(ROOT if os.path.exists(SHARED_ENV) else HERE, ".calls-ledger.jsonl")
PAID_ARMS = ("r", "x")
FREE_ARMS = ("t", "b")
BOOT = os.urandom(6).hex()
PID = os.getpid()
src/app.pythe app
Static files and three JSON endpoints (/api/state, /api/window, /api/triage) on the standard library, port 8792. Renders every window, its alerts and the gold answer with no key configured; only /api/triage calls a provider, and with no API_KEY it returns a calm 200 saying so. Reaches the provider down the same adapter with the same defaults the registered run used -- no reasoning-setting discrepancy, unlike rcv-disc and fin-payrun.
src/app.py
# The minimal local UI. Standard library only -- python -m src.app, then open the printed URL.
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
UI = os.path.join(HERE, "ui")
PORT = int(os.environ.get("PORT", "8792"))
class H(BaseHTTPRequestHandler):
def main():
evals/scoring.pythe scorer
Four things scored, kept apart, never folded into one accuracy number: disposition accuracy, missed_true_positive (the expensive direction for classification), false_correlation (the expensive direction for grouping) and citation_validity. false_correlation is counted over PAIRS rather than groups, because a partial merge inside a bigger group is still a false correlation for every pair it touches and no group-vs-group comparison describes it.
evals/scoring.py
# Score a set of predicted case-window triages against gold. Pure code, shared by
DISPOSITIONS = ("true_positive", "false_positive")
def _gold_group_map(gold_row):
def _pred_group_map(case_groups):
def _norm_kv(s):
def citation_is_real(citation, case_alert_ids, alerts_by_id):
def score(records, gold, windows_by_id, fn_trap_alert_ids, fc_trap_pairs):
evals/baseline.pyfree baseline
A keyword-and-shared-indicator rule written the way a person free-texting a quick triager would write one: trust the three detector types that sound serious, keyword-match the two a human reports, then group anything sharing an entity or ANY raw indicator value and spread the disposition by association. Not weakened for the demonstration -- it is what 'same IP, must be one incident' looks like as code, and it is correct on every window carrying no trap.
evals/baseline.py
# What a simple, dumb rule catches, over the same corpus. Free. No key, no dependency, no model
HERE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
RESULTS = os.path.join(HERE, "results")
TRUSTED_TYPES = ("malware_detection", "data_exfil_alert", "brute_force")
ALARM_RE = re.compile(
def classify(win):
def main():
Start hereThe shortest path into it
src/app.pyrun it and click around. It holds no pipeline logic on purpose; every number on screen comes from a module the eval harness already drives.
tools/build_corpus.pyGenerates 33 case windows and 82 alerts from a fixed seed (SEED=20260819) across five alert types, in five window patterns (7 correlated_tp, 6 uncorrelated_fp, 7 mixed_tp_fp, 5 multi_incident, 8 false_correlation_trap). Gold is DERIVED, never authored beside the data: every alert is tagged at construction time with a disposition and a case_id, and derive_gold() turns those per-alert facts into the alert_dispositions dict and the case_groups list -- used both to write data/gold.jsonl and, independently, by --verify re-reading both files from disk. A swap seam.
src/prompt.pyThe three-part instruction (disposition, correlation, recommendation), the output schema and MAX_TOKENS, declared once and read from here by build() and parse(). States both failure modes to the model as ordinary SOC judgement -- a calm report is not evidence of nothing, a shared address is not evidence of one incident -- rather than as hints about this corpus. A swap seam.
src/triage.pyLoads one case window, calls the model ONCE with every alert in it, and parses the reply into dispositions, case groups and recommendations. One call per WINDOW, never per alert -- that is what makes correlation possible at all, and it is the whole difference from ops-triage's per-window judging, which never compares one alert against another. Never executes containment: there is no function here, in src/app.py or in the UI that locks an account, blocks mail flow or isolates an endpoint.
src/adapters/__init__.pySEAM -- the model. Raw HTTP for every provider (openai-compatible, anthropic), stdlib only. Records finish_reason and reasoning_chars beside the text, which is the only reason the MAX_TOKENS truncation finding was diagnosable without spending again. Takes NO thinking parameter -- see Business.not_good_enough. A swap seam.
src/budget.pyAn append-only ledger written BEFORE each call, shared by every kit under one .env so the daily call cap is on the KEY rather than per kit. Counts calls, not dollars, on purpose.
evals/scoring.pyFour things scored, kept apart, never folded into one accuracy number: disposition accuracy, missed_true_positive (the expensive direction for classification), false_correlation (the expensive direction for grouping) and citation_validity. false_correlation is counted over PAIRS rather than groups, because a partial merge inside a bigger group is still a false correlation for every pair it touches and no group-vs-group comparison describes it.
evals/baseline.pyA keyword-and-shared-indicator rule written the way a person free-texting a quick triager would write one: trust the three detector types that sound serious, keyword-match the two a human reports, then group anything sharing an entity or ANY raw indicator value and spread the disposition by association. Not weakened for the demonstration -- it is what 'same IP, must be one incident' looks like as code, and it is correct on every window carrying no trap.
Every entry above is a file in this kit, listed in the order the pipeline runs it.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
ScenariosWhat it costs at your volume
Everything below is computed from one measured base: 996 input and 1653 output tokens per query, all of it fixed — this kit sends the document's own selected sections, so there is no retrieval depth to turn. Change the volume and the bill scales linearly, because nothing here batches.
Grading really is free here, and that is a property of the ruler rather than a discount
This kit is graded by pure code — the grader makes no model call, needs no key and sends nothing anywhere, so the grading line above is zero at every sampling rate and the leak count stays at nought. What that buys is a ruler you can run on every row, every time, for nothing. What it costs is reach: a code grader can only check what is mechanically checkable.
OperationsFour shapes of workload
The same arithmetic, at volumes worth naming. These are illustrations of scale, not measurements of anyone’s deployment.
Shape
Querys/day
Per day
Per month
Per year
Model follows the controls above.
What this ignoresWhere a real bill diverges
Read before quoting any of it
Output length is assumed constant.A terser or more verbose model moves it, and output is priced above input everywhere.
No volume, committed-use, batch or cache discount is modelled.Each moves real enterprise pricing and none are on a public rate card.
Longer documents cost more and nothing here caps them.This kit selects the document’s own sections in pure Python, so there is no vector store to price — but a corpus of longer records raises the input tokens per query directly, and that is the very number every figure above multiplies.
Grading adds nothing to this bill, and would if you changed the ruler.A code grader is free and local; swapping in an LLM judge would cost more per row than the extraction itself and send every row to a vendor. The instrument is a cost decision, not a footnote.
Rates age.Every one carries the date it was read on the cost page.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
In one lineWhat this kit exposes
This run's alerts are entirely synthetic (tools/build_corpus.py) -- no untrusted party wrote a single description or indicator, and every external IP is from the RFC 5737 documentation ranges. That is a statement about this CORPUS, not about the architecture: in a real deployment, alert descriptions and indicators arrive from detection tools that are themselves reporting attacker-controlled content, which is exactly the shape an indirect prompt injection needs.
Read from the shared .env or the real environment only, never written into the repo, never logged, and scrubbed out of any error string src/app.py returns to the browser.
The experimentWe did not attack it -- and one boundary that should exist does not
An indirect prompt injection needs a field an outside party controls that reaches the prompt verbatim. In a real deployment this kit has several -- every alert description and every indicator VALUE -- and src/prompt.py sends them through unescaped. In THIS corpus every one of them was written by tools/build_corpus.py, so the attack surface exists in the architecture and not in the run. Naming that is the honest version of a resistance rate we do not have. Confirmed by reading the code, not by a run, on 2026-08-20 -- no attack run was fired; see redteam.why.
Boundary checked
A naive build
This kit
The kit never executes a containment action
an agent wired to a containment API locks the account it just decided was compromised
holds in CODE, not just in the prompt: there is no function in src/triage.py, src/app.py or the UI that locks an account, blocks mail flow or isolates an endpoint. 0 of 33 calls wrote anything anywhere.
Alert text reaches the model as data, in a labelled block
alert descriptions concatenated into the instruction text with no separation
the alert block is assembled by src/prompt.py::_alert_block() and sent in the USER message; the instruction lives in SYSTEM. That is a separation, not a defence -- no escaping, no delimiter check, and it has never been attacked.
A citation must name an indicator that really exists on an alert in that case
a plausible-sounding indicator string is rendered as evidence
checked OFFLINE by evals/scoring.py::citation_is_real -- exact (name, value) lookup, 109 of 109 valid this run. NOT checked live: src/app.py renders whatever the model cited with no verification step.
The gold answers never reach the prompt
the labels are loaded alongside the data and leak into context
holds in code: src/triage.py::load_gold() is never called by check(), and the docstring says so in as many words. The app shows gold to the READER on purpose -- it is a demo of the mechanic, not a blind quiz -- but it is not in the prompt.
Each boundary above was checked by reading the call sites, not by an attack trial. Gate 1 is the strong one -- it holds because the function does not exist, which is the only kind of guarantee this kit offers. Gates 2 and 3 are separations and offline checks respectively, and neither is enforced live.
The result0 attack trials, one boundary that DOES hold in code (nothing here can execute containment), and one that does not (nothing cross-checks the model's own citations or groupings before the UI renders them).
2externally-authored field types a live deployment would carry (every alert description and every indicator value) -- synthetic on this run's corpus
0 of 0attack trials run
0 of 109fabricated citations -- measured over the run, not under attack
This run's corpus is entirely generated -- no alert's description or indicator was written by an outside party -- so there was no live untrusted input to attack with and no attack was attempted. The 0 fabricated citations is a MEASURED result over 109 citations, not an attack result.
Read this twice
This kit has no code-level check on its own model's output. Nothing verifies that a cited indicator exists before the UI renders it, that an alert placed in a case was also called true_positive, or that the case groups partition the true positives at all. All three are checked OFFLINE by evals/scoring.py against gold, which a real deployment does not have. The 100% citation validity on this page is a measurement of one run, not a guarantee about the next one.
HonestyWhat this does not prove
Whether the prompt-level boundaries survive hostile alert text -- no red-team run exists for this kit.
Whether a live citation check would ever have caught a real fabrication -- this run's model never produced one to test against.
Whether a model that grouped an alert it also called false_positive would be noticed -- nothing in code checks that consistency, live or offline.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
In one lineThe only guardrail is the prompt
One instruction, no enforcement. It mostly holds — and mostly is the whole subject.
the guardrail, verbatim
Two alerts belong in the same case only when the evidence connects them causally -- a shared entity, a close timestamp or one matching indicator value is not a link on its own. And a calm or routine-sounding report is not evidence of a false positive.
src/prompt.py -- SYSTEM, stated as ordinary SOC judgement the model must apply, not as a hint about this corpus. Prompt-level ONLY: nothing in src/triage.py or src/app.py re-checks a grouping, a disposition or a citation before the UI renders them.
EvidenceDoes it hold?
What
Measured
No code path in this kit executes a containment action
0 of 33 calls in r001-it-sectriage resulted in any write or any outbound action -- src/triage.py::check() and src/app.py's /api/triage both only ever return an answer. The guarantee is that the function does not exist.
The gold answers never enter the prompt
src/triage.py::load_gold() is never called by check(); confirmed by reading both. The app shows gold to the reader deliberately, but never to the model.
The prompt-level correlation rule held on this run
0 of 47 gold-different pairs merged, including 0 of the 8 planted coincidental-indicator pairs the free floor merges 8 of 8 times -- but see fails for what does, and does not, enforce this.
The prompt-level mundane-phishing rule held on this run
0 of 52 gold true positives called false_positive, including 0 of the 3 planted mundane-worded reports.
The limitWhat a guardrail is not
IT IS NOT A CODE-ENFORCED CHECK. Both rules live only in src/prompt.py's SYSTEM text -- nothing in src/triage.py or src/app.py recomputes them. See fails.
It does not verify that a citation is relevant, only (offline, in the eval) that it exists.
It does not make the disposition or the grouping correct -- see Eval.taxonomy for what this run measured, not what any guardrail guarantees.
It is not a defence against hostile or fabricated alert text -- no red-team run exists for this kit (see the security page).
It is not a correlation ENGINE. This kit judges the bundle it is handed; which alerts get bundled together is decided upstream and is not guarded here at all.
WatchedWhat is watched, and why that one
1run recorded
0 / 0deterministic metrics exact
+0.0%largest move — latency
0band breaches
—model half · needs provider URL
The kit measures 18 watched figures plus 8 comparability guards and 5 reference constants — three duties, not one list. 0 of the watched are grader rates, each on its own grader page; this board watches the ones that decide when to look again. Each square is one watched figure; filled means the latest run measured it.
15 measured by the latest run3 need the model half
Metric
Owner
Role
Why this one
disposition-correlation-and-citation-exact-match
Disposition exact match, pair-wise false correlation, and citation-existence checks
alarm
false_correlation on the trap subset as a raw count, never folded into disposition accuracy -- it is the expensive-direction error this kit's second trap exists to catch; missed_true_positive on the trap subset specifically, since a keyword reader fails exactly there and nowhere else (see baseline_note); answered vs asked -- 100% this run, but a truncated reply answers fewer alerts while every alert it DID answer stays correct, which is precisely how the 3000-token run manufactured a 34.6% miss rate — alarm on Any nonzero false_correlation on the trap subset, or any nonzero missed_true_positive on the trap subset, on any run. Both zero this run.
GuardsBefore any comparison
Preconditions, not alarms. If one moves, the runs were not measuring the same system and nothing else on this board may be differenced.
Guard
Value
If it moves
corpus.doc_count
33
different corpus — nothing is comparable
corpus.bytes
33,449
case windows edited — the count held, the bytes did not
split.count
33
the characters per window's assembled alert block count moved — a different set was scored
split.size_p50
656
the median size of one characters per window's assembled alert block moved
split.size_p95
893
the 95th-percentile size of one characters per window's assembled alert block moved
dataset.rows
33
the test set changed — every rate has a new denominator
tokens.context_limit
1,000,000
different model family
index.build_seconds
0.0
no index is built here — a change is in corpus preparation, not an index rebuild
The reference constants — the null-grader baseline, the adjudication matrix, the negative count — are the third duty. They are never re-measured per run: a baseline that moves with the thing it measures is not a baseline.
BandsThe alarm bands
Derived from evidence, never picked: exact-match where two runs reproduced a metric to the digit, wider than the measured same-input drift for latency, one row of the negative class for any rate — denominator printed beside it.
Group
Band
Denominator
Where the band comes from
disposition accuracy
100 -- 100% this run
82 answered alerts
measured, one run
alerts answered
not yet known
82 alerts
100% this run -- every alert drew a disposition, none silently dropped. At MAX_TOKENS=3000 this same figure was 74.4%, which is how the truncation was visible in the record all along. No repeat at 8192 -- r001 ran once.
missed true positive
0 -- zero this run, one of this kit's two headline guardrail metrics
52 gold true-positive alerts
measured, one run
missed true positive, planted trap subset
0 -- zero this run; the free floor misses 2 of these 3
3 planted mundane-phishing alerts
measured, one run
false correlation
0 -- zero this run, this kit's other headline guardrail metric
47 gold-different alert pairs
measured, one run
false correlation, planted trap subset
0 -- zero this run; the free floor merges 8 of these 8
8 planted coincidental-indicator pairs
measured, one run
citation validity
100 -- 100% this run
109 citations across 35 drafted recommendations
measured, one run
latency
not yet known
33 calls
p50 9,758 ms, p95 47,346 ms on r001-it-sectriage -- reasoning left on, with no way to turn it off. One recorded run, not a distribution. The p95 is one window, cw026.
input volume
0 -- fixed by the corpus and the prompt, not the model
33 calls
32,892 input tokens on r001-it-sectriage. Any movement means the prompt or the corpus changed.
output volume
not yet known
33 calls
54,550 output tokens on r001-it-sectriage -- model-specific, and an estimated 95.7% of it is the provider's reasoning pass. Per-call it ranged 105 to 5,442.
reasoning share of output
not yet known
33 calls
an estimated 95.7% this run (about 52,183 of 54,550 output tokens, from 208,731 recorded reasoning characters at roughly four characters per token), one run, one setting -- and this kit ships no way to run the other setting. Narrative only, not a registered metric key: the run record carries reasoning_chars per call but no reasoning-share field.
HistoryRun history
1 recorded run. A dash is a metric that run did not measure, and a dash is never differenced: not measured is a third state, neither a value nor a zero, and reading it as zero is how a board reports an improvement on a run that never looked.
Metric
r001-it-sectriage 2026-08-20
answered, %
100.0
citation validity, %
100.0
disposition accuracy, %
100.0
false correlation
0
false correlation rate, %
0.0
false correlation trap
0
false correlation trap rate, %
0.0
input tokens, whole run
32892
model latency p50 ms
9758.00
model latency p95 ms
47346.00
missed true positive
0
missed true positive rate, %
0.0
missed true positive trap
0
missed true positive trap rate, %
0.0
output tokens, whole run
54550
one run A verdict is a comparison and one run has nothing to compare with, so the Verdict column is left off rather than filled with 15 chips that all say so.
DeviationsWhat deviated
0 breaches across 1 runs
The day a band breaks, a row renders here carrying the metric, the two runs, the band it crossed and what changed between them — model, prompt, chunker or top_k, read off the run records’ guards. A deviation with no diff beside it is an alarm nobody can act on.
RippleThe ripple map
Which figures move together, so one alarm reads as a system. Every edge carries its basis: measured names the run that showed it; reasoning admits nothing has varied it yet.
Lever
What moves
Basis
Evidence
the reply-token ceiling (MAX_TOKENS in src/prompt.py)
missed_true_positive: 34.6% at 3000 -> an inflated rate at 4096 -> 0.0% at 8192, on the identical corpus with the identical scorer. None of the movement was the model changing its mind; it was replies being cut off
measured
7 of 33 calls finish_reason='length' at 3000 with only 61 of 82 alerts answered and 100% accuracy on those; 5 of 33 at 4096; 0 of 33 at 8192 on the published run. Recorded in src/prompt.py's MAX_TOKENS comment and in the kit README.
the planted coincidental indicator (8 of the 26 windows presenting a merge-or-don't decision)
false_correlation on the 8 trap pairs: 100% (keyword-and-shared-indicator floor, all 8 merged) -> 0% (the fast tier) on the identical windows
measured
results/eval-b000-keyword-rule.json vs results/eval-r001-it-sectriage.json, same 33 windows, one variable (which decider reads them).
whether the provider's reasoning pass runs
not measured -- this kit's adapter has no thinking parameter, so the reasoning-off configuration has never been run
reasoning
every one of the 33 records carries a nonzero reasoning_chars (331-23,475), and the run file has no top-level thinking field because nothing ever sent one. See Business.not_good_enough.
FiresWhat fires when a band breaks
Where evals stop sitting beside the guardrail and start feeding it. Where the answer is nothing automatic, the row says so — hiding it would undo “What a guardrail is not” above.
Band
What fires
disposition accuracy
any drop below 100% on a re-run
alerts answered
anything below 100%, on any run -- it means a reply was cut off
missed true positive
any nonzero value, on any run
missed true positive, planted trap subset
any nonzero value, on any run
false correlation
any nonzero value, on any run
false correlation, planted trap subset
any nonzero value, on any run
citation validity
any drop below 100% on a re-run -- a fabricated citation is a fabricated fact
latency
nothing yet
input volume
any change without a corresponding change to prompt or corpus
output volume
any call reaching MAX_TOKENS (8192) -- that is a truncated answer, not a long one
reasoning share of output
nothing yet
NextThe three you would add first
A live, code-level consistency check: every alert in a case_group must also carry a true_positive disposition, the groups must not overlap, and every citation must resolve to a real (indicator, value) on an alert in that case -- before the UI renders any of it.citation_is_real already exists in evals/scoring.py and needs no gold to run; it is checked against the eval set and not against the live reply, which is the whole gap.
Surface finish_reason in /api/triage and refuse to score a truncated reply as an answer.This kit has already published a fake 34.6% failure rate from exactly this, twice over. A truncated reply is a reliability failure, not a cautious judgement, and the two must never be counted as the same thing.
A red-team run against the alert description and indicator fields.In a real deployment those fields quote attacker-controlled material and this kit's architecture treats them as trusted with no escaping -- unmeasured (see the security page's posture).
None of these is built here. They are named against the seam they sit on so they can be argued with.
CadenceWhen to re-run, and what it costs
Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS, and read finish_reason on every record before believing any rate the re-run reports.
What this cannot tell you
Whether the same 100%/0% figures hold on a second run, or on a window wider than 4 alerts.
Whether a live, code-level consistency check (see add_first) would ever have caught a real disagreement -- this run's model never produced one.
Whether the prompt-only correlation and grounding rules hold against hostile alert text -- no red-team run exists for this kit.
Whether the answers change with reasoning disabled -- this kit cannot currently run that configuration.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
In one lineNo framework, deliberately
No framework and no dependencies beyond the standard library -- requirements.txt is a comment-only file explaining that the emptiness is load-bearing. The whole triage decision is three files: src/prompt.py, src/triage.py and src/adapters/__init__.py.
The mappingSeam by seam, who owns what
Seam
File here
Framework equivalent
Note
the corpus
tools/build_corpus.py
none -- a seeded generator
33 case windows and 82 alerts generated from a fixed seed, never fetched. Deriving gold from each alert's own construction-time disposition and case_id (never the pattern that seeded the window) via derive_gold() is what keeps the internal-consistency check honest -- see data/SOURCES.md.
prompt assembly
src/prompt.py
prompt templates
the three-part instruction, the output schema and MAX_TOKENS are one declaration and SYSTEM is built from it -- so the prompt can be published verbatim, which a template assembled two calls away cannot be.
the model
src/adapters/__init__.py
chat model wrappers / vendor SDKs
raw HTTP over urllib for every provider, so a forker runs this on whichever key they hold. It records finish_reason and reasoning_chars, which is the only reason this kit's truncation defect was diagnosable without paying for a fourth run -- and it is also where the missing thinking parameter would go.
evaluation
evals/scoring.py
eval harnesses
exact match over a two-value disposition plus a pair-wise grouping comparison and an indicator lookup is a few counters and a union-find-free double loop, not a platform.
spend control
src/budget.py
none
an append-only ledger written BEFORE each call, shared by every kit under one .env so the cap is on the KEY rather than per kit.
The seams are the ones this kit already publishes, so the mapping cannot drift from the code.
GraphsWhere LangGraph changes the answer
Nothing here is a graph. The pipeline has one path per window -- load, assemble, prompt, call, parse -- with no branching and no state carried between windows. A framework would add an orchestrator to a for-loop.
The other sideWhat a framework costs you
A third disposition value this kit's fixed two-value vocabulary does not name (a 'needs enrichment' state, say) needs a hand-written change in src/prompt.py plus new corpus logic in tools/build_corpus.py, rather than being configured declaratively.
No built-in retry/backoff beyond src/adapters.py's own bounded retry -- a framework's queue and worker model is not here, so a real deployment adds its own scheduling.
No built-in observability beyond what evals/run.py prints and writes to results/ -- a framework's tracing/dashboard integration is not here. What IS recorded, and mattered, is finish_reason and reasoning_chars per call.
No token-budget management. A framework that sized max_tokens against observed output would have caught the truncation this kit had to find twice by hand; the stdlib-only design has one constant and a comment.
No built-in output validation beyond src/prompt.py::parse()'s tolerant-but-not-creative JSON parse -- a framework with a schema-and-consistency layer might catch a model that groups an alert it also called false_positive before it renders; this kit does not check that in code (see Guardrails.fails).
What we could NOT verify
No port to any framework was actually built and timed. The costs above are read off this kit's own code and the frameworks' documented feature sets, not measured against a working port.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
In one lineWhat one run of this kit actually cost, in time and tokens
Measured on r001-it-sectriage on the fast tier, 2026-08-20. This kit records telemetry measured per run — 4 of the 6 readings on this axis, 4 of them with a band the guardrail board can judge against.
The readingWhat the last run measured
Reading
Last run
Band
What fires
Model, median
9,758 ms
not yet known
nothing yet
Model, p95
47,346 ms
not yet known
nothing yet
Input tokens
32,892
0 -- fixed by the corpus and the prompt, not the model
any change without a corresponding change to prompt or corpus
Output tokens
54,550
not yet known
any call reaching MAX_TOKENS (8192) -- that is a truncated answer, not a long one
No movement column. This is the only run on record, so there is nothing to move against. Said once here rather than as a dash on every row.
Retrieval not shown. No run of this kit has ever recorded a retrieval timing, because the pipeline has no such stage. Those readings are left out rather than dashed: a dash says “this run did not measure it”, which promises a later run could.
The other sideWhat this does not tell you
What we could NOT verify
No stage-by-stage span tree. These are whole-call timings; which part of the pipeline spent the time is not recorded by this kit's harness.
Latency is a run-level median and p95. The spread between the fastest and slowest single call is not stored, so the distribution cannot be drawn.
No earlier run is comparable to this one (a different model tier or a changed guard), so nothing on this page is a trend yet.
One machine, one network, one region. Latency includes the provider round trip from a single location and will not transfer to yours.
The first of those is the one worth doing something about, and it is a solved problem: instrumenting a pipeline so each step reports its own duration is a couple of lines. Observability / tracing has the code, and what the result looks like once it is running.
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
In one lineWhat you provision, honestly
Everything measured on this page says where it was measured; everything not measured says that — the same contract as the other thirteen pages. Where a retrieval index would be on other kits, this one has: no index — windows are cut from the log stream per run.
The machineWhat this needs
Dependencies
nothing beyond the standard library — requirements.txt names no package
Configuration
7 environment variables — the full knob table is at the bottom of this page
Last verified run
2026-08-20, across 1 committed record
Enumerated from the kit’s own source by build/facts/envscan.py, committed, and drift-checked — never typed into this page.
Your dataWhere everything sits
Artifact
Where
Leaves your machine?
case windows
data/windows.jsonl — 33 windows, 82 alerts, fixed seed, no clock read; your disk
one window at a time, whole, per call. ⚠︎ The heading above is the triage flow's own locked wording: on THIS kit nothing is cut from a stream — the 2-4 alert bundles arrive already formed and src/triage.py judges what it is handed. See the windowing rung below.
gold dispositions and case groups
data/gold.jsonl — 33 rows, each naming its trap kind and the exact alert ids that carry it
never — scoring is in-process in evals/scoring.py, no judge model
the key
.env at the repo root, shared by every kit — never committed
only inside the Authorization header
The wireWhat crosses it, and when
Runtime call site
Destination
src/adapters/__init__.py line 55
the configured BASE_URL
Build-time (corpus tooling, run once)
Destination
none — this kit’s corpus is generated locally
Zero-egress topology: point BASE_URL at a local OpenAI-compatible server and nothing leaves the machine — statically true of the shipped adapter, whose only credential travels in the Authorization header.
The third state
What the provider retains, trains on or logs once a request arrives is not measured here — it is provider-dependent, and this page does not assert what no run of ours can verify.
The shapeHow it ships, and what you add
The key
Read from the shared .env or the real environment only, never written into the repo, never logged, and scrubbed out of any error string src/app.py returns to the browser.
You add
your reverse proxy and your identity in front of one process — the kit deliberately ships neither, because identity belongs to your estate, not to a kit
Never ships
Docker, Helm, a queue, a database. The product is a folder of readable Python; what to run it under is this page’s job to state, not the kit’s job to impose
The ladderEach decision, and what you provision past its ceiling
Decision
The kit ships
Measured here
Past the ceiling
What stops being valid
windowing
the windows as authored — tools/build_corpus.py bundles 2-4 alerts per candidate group and src/triage.py judges whatever it is handed; nothing in this kit cuts a stream
nothing. This is the one decision this kit does NOT make and does NOT measure — 18 of 33 windows carry 2 alerts, 14 carry 3, exactly 1 carries 4, and no other width was run (data/SOURCES.md and lenses.Architecture.breaks_at_scale; no re-cut experiment exists)
a real deployment's own entity/time correlation window governs this entirely. Too loose over-merges unrelated alerts before a model ever sees them; too tight splits one real incident across two windows that are never compared
every published score, because the pairwise correlation judgement is a property of the bundle. A wider window is more merge-or-don't decisions per call against the same reasoning budget, and it is untested
model
one HTTP completion call per window behind src/adapters/__init__.py — any OpenAI-compatible endpoint or Anthropic; one JSON object back with dispositions, case groups and recommendations
one tier, one run: 100% disposition accuracy, 0 missed true positives, 0 false correlations, p50 9,758 ms, p95 47,346 ms (lenses.Eval.scores and lenses.Cost.cost_by_model, r001-it-sectriage)
a hosted provider or a local server — the .env decides, not the code. Two things to size before you switch: MAX_TOKENS must clear your model's reasoning appetite (8192 here, after 3000 and 4096 both truncated), and whether the model reasons by default decides most of your bill
published accuracy, latency and cost are all per-model — the floor and the scorer do not move, so the comparison re-runs on yours for the price of 33 calls
labels
data/gold.jsonl — 33 windows whose dispositions and case groups are DERIVED by derive_gold() from each alert's construction-time facts; tools/build_corpus.py --verify re-derives every row against the files on disk and refuses the set on disagreement
zero drift on --verify; no false positive appears in any case group, and every planted trap_pair is genuinely gold-separate (data/SOURCES.md and lenses.Eval.validated)
label windows from your own queue — and expect the hard part to be the grouping, not the disposition: deciding that two alerts were the same incident is a judgement a real SOC makes after the fact, with information the alerts themselves never carried
the labels are a property of the generator AND the bundling; your stream has no planted truth, so the labelled set is yours to build before any score means anything
Every threshold is a measurement with its provenance beside it, or says is_measured:false. A ceiling with no seam says so — it is the point the kit is outgrown, never a forced join.
When it failsFailure signatures
What you see
What it is
First move
two alerts on different users merged into one case because they share a source IP
the correlation trap this kit's second failure mode is built from — most often a shared building egress or VPN concentrator address, not one attacker
read what the shared value actually IS before treating it as a link, and score correlation over PAIRS rather than groups so a partial merge inside a bigger group cannot hide (lenses.Eval.taxonomy (false_correlation_trap) and example_row, r001-it-sectriage)
a calm 'is this legit?' phishing report waved through as a false positive
the first planted trap — a genuine report worded like routine correspondence, which a keyword reader has nothing to catch; the free floor misses 2 of the 3
judge the sender domain, the link and the ask, not the tone — and check whether a loud neighbour in the same window is evidence for the quiet one, as it is in cw006 (lenses.Eval.baseline_note and data/SOURCES.md, b000-keyword-rule)
No machine symptom — this failure leaves no trace in any output.
a reply that ran out of tokens looks EXACTLY like a model that missed an incident. At MAX_TOKENS=3000 this kit scored a 34.6% missed-true-positive rate that was entirely truncation — 7 of 33 calls cut off, only 61 of 82 alerts answered, and 100% accuracy on every alert that did get an answer. The control is recording finish_reason and reasoning_chars on every call and reading them BEFORE believing a failure rate (lenses.Business.not_good_enough)
Concurrency and GPU sizing — one serial call per window, nothing measured past 33. Provider-side retention — provider-dependent, and on this kit the payload is alert text naming users and hosts, so that unknown IS the posture question. Where a window stops fitting the reasoning budget: this corpus's widest window is 4 alerts and its worst call spent 5,442 of 8,192 output tokens, so the ceiling was approached but never found. The reasoning-off configuration, which this kit cannot currently produce. And what a long incident costs: there is no cross-window deduplication and a run that dies at window 20 starts again — no checkpointing, no resumption.
The corpus licence, from the Data lens: MIT (this repository's own licence) Your corpus’s licence is yours to verify, and the labels you write are about your documents.
Disposition exact match, pair-wise false correlation, and citation-existence checks
Sort a security team's alert bundles into real incidents
PresenterOpens the private repo. Visible to admins only.
In one lineDisposition exact match, pair-wise false correlation, and citation-existence checks
Does each alert's true_positive/false_positive call match gold? Does the model avoid calling a gold true positive a false positive (missed_true_positive), overall and on the 3 deliberately mundane-worded phishing reports? Does it avoid putting two alerts in one case that gold keeps apart (false_correlation), overall and on the 8 planted coincidental-indicator pairs? Does every citation name an indicator that actually exists on an alert in that case?
$0.00per 1,000 case windows
nodata leaves your network
yessame answer every time
MethodHow the test was run
evals/scoring.py::score, in-process, no key and no model. The same function evals/baseline.py and evals/run.py both call -- a baseline and a real run scored by two different scorers cannot be compared honestly.
The inputOne real row, seen by every grader
Case window
cw026
Window pattern
false_correlation_trap
The coincidence
source_ip=198.51.100.177 -- on all three alerts
The alerts in the window
ALT-0062 suspicious_login (user bcollins, 22:02) -- successful login right after a string of failed attempts from the same source IP. ALT-0061 brute_force (user bcollins, 22:08) -- 18 failed attempts against one account from a single source IP, followed by a successful login. ALT-0063 suspicious_login (user cbennett, 22:14) -- login from an unfamiliar location and new device; the user confirmed by phone they are travelling for a conference this week.
[ALT-0061, ALT-0062] -- one incident; ALT-0063 belongs to no case
Keyword-and-shared-indicator floor's call
all three true_positive, all three in ONE case -- a false correlation on the planted pair, and a benign travelling user swept into an incident
The model's call
ALT-0061 and ALT-0062 true_positive and grouped together; ALT-0063 false_positive and in no case -- correct on every axis
The model's drafted recommendation
Lock target_account=bcollins and require a password reset; verify and revoke any active sessions associated with this login before approving.
Scored as
correct -- the trap the free floor fails and the fast tier does not. It is also the run's single most expensive call: 5,442 output tokens, 22,431 reasoning characters and 47,346 ms, which IS this kit's published p95 latency.
Grader
Verdict
Why
Disposition exact match, pair-wise false correlation, and citation-existence checks
correct
cw026: three alerts all sharing source_ip=198.51.100.177, two of them also sharing device_id=dev-c2277. ALT-0061 (brute_force, user bcollins) and ALT-0062 (the successful login that followed it, same user) are one real incident; ALT-0063 is a different user, cbennett, logging in from an unfamiliar location on a new device, who confirmed by phone they are travelling. Gold: two true positives in one case, one false positive in no case. The keyword-and-shared-indicator floor grouped all three and called all three true positives -- a false correlation on the planted pair AND a wrong disposition. The fast tier returned exactly gold, and cited target_account=bcollins and attempt_count=18 on its lockout recommendation, both real indicators on ALT-0061.
The formulaWhat it computes
disposition_accuracy = correct dispositions / answered, over answered=82. missed_true_positive = a gold true_positive the model called false_positive, over the 52 gold true positives; its trap subset uses exactly the 3 alert ids gold's trap_alert_ids names. false_correlation is counted over PAIRS: a gold-different pair is any two alerts not in the same true-positive case together in gold (including a true positive paired with a false positive), giving 47 pairs; its trap subset is exactly the one planted trap_pair per false_correlation window, 8 of them. citation_validity = citations whose (name, value) exists on an alert in that case / all 109 citations drafted.
The analysisWhat it actually did
Model
Result
the fast tier
100.0% disposition accuracy · 2 more measured on this row
In operationWhat to monitor
Reference standard: tools/build_corpus.py's derive_gold(), which turns each alert's construction-time disposition and case_id into the window's alert_dispositions and case_groups -- re-derived independently via --verify against the windows and gold actually written to disk, asserted zero drift.
These rates are UNKNOWN, on purpose
This grader's own error rate is not separately measured -- it IS the reference. What can go wrong is the corpus's own trap design, which comes from tools/build_corpus.py's fixed templates.
Watch these
false_correlation on the trap subset as a raw count, never folded into disposition accuracy -- it is the expensive-direction error this kit's second trap exists to catch
missed_true_positive on the trap subset specifically, since a keyword reader fails exactly there and nowhere else (see baseline_note)
answered vs asked -- 100% this run, but a truncated reply answers fewer alerts while every alert it DID answer stays correct, which is precisely how the 3000-token run manufactured a 34.6% miss rate
Alarm on
Any nonzero false_correlation on the trap subset, or any nonzero missed_true_positive on the trap subset, on any run. Both zero this run.
How tight can the band be? There is no tolerance band -- every field is an exact match (a two-value disposition, a grouping compared pair by pair, an indicator string looked up exactly after normalising 'k=v'/'k: v' spacing and case). Nothing here is a continuous quantity to round.
Cadence: Re-score on any change to src/prompt.py or tools/build_corpus.py -- the first changes what is asked, the second changes what is asked ABOUT. Re-run evals/run.py (paid) on any change to MAX_TOKENS, which this kit has already had to do twice.
The decisionWhen to reach for it
Use it
Gold is DERIVED from the same generated alert records the model reads -- true of every kit corpus, never true of a real SOC's own queue.
Do not use it
The truth is not known in advance -- the normal state of a real alert queue, and the reason this corpus is generated rather than captured.
A living map of modern AI — kept current every morning