The ideaThree seats, one pipeline
An enterprise does not have an AI dashboard. It has three, and they disagree. A developer is measured on what reaches production and whether it stays there. An architect is measured on what that traffic costs and whether it stayed inside policy. A project manager is measured on whether any of it turned into delivered value.
Most dashboards pick one seat and quietly blame the other two. Spend is up reads as an architecture failure until you notice deploys doubled. Cycle time is long reads as a developer problem until you notice the queue is upstream of them. Put the three scorecards side by side and a number can be read against the thing that actually moved it.
At a glanceHow the three scorecards chain together
Each seat’s output is the next seat’s input. Deploy volume sets the traffic an architect has to route affordably; routing quality sets the latency and spend a project manager reports as value. That is why a failure in one seat surfaces as a bad number two seats downstream, and why reading any one scorecard alone tends to blame the wrong person.
The scorecardWhat each seat is judged on
One system, three readings. Change the monthly budget and the credit figures recompute against it.
Did it ship, and did it stay shipped
| Metric | Worked example | What a bad number actually means |
|---|---|---|
| Deploys reaching production | 142 /mo | Low means batching — big, rare, risky releases instead of small safe ones. The fix is upstream of the developer. |
| Go-live success rate | 94.4% | Low means the tests pass and reality does not. Look for what the rollbacks have in common before you look at who wrote them. |
| Credit burn against budget | 21,380 / 30,000 | Rising without more output is agent sprawl, not more work getting done. Check the ratio to deploys, never the total on its own. |
| Rework rate | 6.8% | High means the specification was unclear before the agent started. It is an acceptance-criteria problem wearing a code-quality costume. |
What the traffic costs, and whether it stayed in policy
| Metric | Worked example | What a bad number actually means |
|---|---|---|
| Token cost saved by routing | 38% | Low means frontier prices are being paid for work a small model handles. This is usually the largest single saving available and the least often taken. |
| Small-model / frontier mix | 61% / 39% | The lever behind the saving above. Move this before negotiating rates — routing beats discounting. |
| Guardrail compliance | 99.2% | Anything under 100% needs a named, dated exception with an owner. A standing 99% is not a rounding error, it is an unreviewed decision. |
| Uptime and p95 latency | 99.95% · 1.9s | Where routing savings get paid back. Cheap and slow is not a win, and p95 is the number users feel — the average hides them. |
Did any of it become delivered value
| Metric | Worked example | What a bad number actually means |
|---|---|---|
| Cycle time, idea to production | 11 days | Long usually means the queue, not the coding. Measure the waiting, not the working, or you will optimise the fastest part of the process. |
| Credit ROI | 1.9× per credit | Ties spend to outcome instead of to activity. Activity is not delivery, and a busy pipeline can return nothing. |
| Eval scores, relevance and accuracy | 4.4 / 5 · 91% | Falling means the thing you already shipped got worse while every other number looked fine. It is the only metric here that watches the product itself. |
| Delivery predictability | 8 of 10 committed | Low means the estimates are wrong, not that the team is slow. Those are different problems with opposite fixes. |
Illustrative numbers throughout — these are shapes, not measurements. Nothing on this page is read from a live system. Swap in your own figures; the structure is the point.
MethodHow to read these numbers
- Never read one seat alone.Every metric here has a plausible innocent explanation sitting in another column. Credit burn rising alongside deploys is throughput; rising without them is sprawl.
- Separate leading from lagging.Cycle time and rework move first. Eval scores and credit ROI move last, which is why a team can feel fine for a month after it has stopped being fine.
- Ratios over totals.A total says how busy you were. Credits per deploy, exceptions per thousand requests and value per credit say whether it was worth it.
- A target band, not a target.100% guardrail compliance with no exception process means nobody is asking; 94% go-live success may be correct for a team shipping small and often.
- Name the owner of each number.A metric with three owners has none, and every dispute on this page is really a dispute about who is accountable for it.