Home › Operate › Enterprise AI Operations
🎛️ Operate

Enterprise AI Operations

One delivery system, three seats — and the numbers each one is actually judged on.

The ideaThree seats, one pipeline

An enterprise does not have an AI dashboard. It has three, and they disagree. A developer is measured on what reaches production and whether it stays there. An architect is measured on what that traffic costs and whether it stayed inside policy. A project manager is measured on whether any of it turned into delivered value.

Most dashboards pick one seat and quietly blame the other two. Spend is up reads as an architecture failure until you notice deploys doubled. Cycle time is long reads as a developer problem until you notice the queue is upstream of them. Put the three scorecards side by side and a number can be read against the thing that actually moved it.

At a glanceHow the three scorecards chain together

Developer throughput feeds architect cost and safety, which feeds project-manager delivered value

Each seat’s output is the next seat’s input. Deploy volume sets the traffic an architect has to route affordably; routing quality sets the latency and spend a project manager reports as value. That is why a failure in one seat surfaces as a bad number two seats downstream, and why reading any one scorecard alone tends to blame the wrong person.

The scorecardWhat each seat is judged on

One system, three readings. Change the monthly budget and the credit figures recompute against it.

A credit is whatever unit your organisation budgets AI work in — tokens, currency, seats or vendor credits. 30,000 is a worked example, not a recommendation.

Did it ship, and did it stay shipped

142
deploys reaching production
+24 vs last month
94.4%
go-live success, no rollback within 24h
+2.1 pts
21,380
of 30,000 monthly credits
6.8%
rework rate, merged then reopened
1.4 pts better
MetricWorked exampleWhat a bad number actually means
Deploys reaching production142 /moLow means batching — big, rare, risky releases instead of small safe ones. The fix is upstream of the developer.
Go-live success rate94.4%Low means the tests pass and reality does not. Look for what the rollbacks have in common before you look at who wrote them.
Credit burn against budget21,380 / 30,000Rising without more output is agent sprawl, not more work getting done. Check the ratio to deploys, never the total on its own.
Rework rate6.8%High means the specification was unclear before the agent started. It is an acceptance-criteria problem wearing a code-quality costume.

What the traffic costs, and whether it stayed in policy

38%
token cost saved by routing
+6 pts
61 / 39
small-model / frontier traffic split
99.2%
requests passing guardrail policy
8 exceptions open
1.9s
p95 latency · 99.95% uptime
within band
MetricWorked exampleWhat a bad number actually means
Token cost saved by routing38%Low means frontier prices are being paid for work a small model handles. This is usually the largest single saving available and the least often taken.
Small-model / frontier mix61% / 39%The lever behind the saving above. Move this before negotiating rates — routing beats discounting.
Guardrail compliance99.2%Anything under 100% needs a named, dated exception with an owner. A standing 99% is not a rounding error, it is an unreviewed decision.
Uptime and p95 latency99.95% · 1.9sWhere routing savings get paid back. Cheap and slow is not a win, and p95 is the number users feel — the average hides them.

Did any of it become delivered value

11 days
cycle time, idea to production
3 days faster
1.9×
value returned per credit spent
4.4 / 5
eval relevance · 91% accuracy
holding
8 of 10
committed items actually shipped
was 9
MetricWorked exampleWhat a bad number actually means
Cycle time, idea to production11 daysLong usually means the queue, not the coding. Measure the waiting, not the working, or you will optimise the fastest part of the process.
Credit ROI1.9× per creditTies spend to outcome instead of to activity. Activity is not delivery, and a busy pipeline can return nothing.
Eval scores, relevance and accuracy4.4 / 5 · 91%Falling means the thing you already shipped got worse while every other number looked fine. It is the only metric here that watches the product itself.
Delivery predictability8 of 10 committedLow means the estimates are wrong, not that the team is slow. Those are different problems with opposite fixes.

Illustrative numbers throughout — these are shapes, not measurements. Nothing on this page is read from a live system. Swap in your own figures; the structure is the point.

MethodHow to read these numbers

  • Never read one seat alone.Every metric here has a plausible innocent explanation sitting in another column. Credit burn rising alongside deploys is throughput; rising without them is sprawl.
  • Separate leading from lagging.Cycle time and rework move first. Eval scores and credit ROI move last, which is why a team can feel fine for a month after it has stopped being fine.
  • Ratios over totals.A total says how busy you were. Credits per deploy, exceptions per thousand requests and value per credit say whether it was worth it.
  • A target band, not a target.100% guardrail compliance with no exception process means nobody is asking; 94% go-live success may be correct for a team shipping small and often.
  • Name the owner of each number.A metric with three owners has none, and every dispute on this page is really a dispute about who is accountable for it.
A living map of modern AI — kept current every morning