
May 28, 2026
in /
Deployment
4 min read
What to measure once the automation is actually running
Model accuracy stops being the interesting number the day real work arrives. What matters after that is what the operation does differently, and who can see it.
Nukes AI
AI automation team
Accuracy is the number that gets an automation approved and the wrong number to run it on. Once real work is flowing, what matters is whether the operation changed shape: how many cases still need a person, how quickly they get one, and whether anybody outside the build team can see the answer without asking an engineer for it.
Accuracy stops being useful
A validation score describes how a model behaved on data collected before it existed. The day it starts making decisions the inputs change, because the process changed around it, and the score goes on describing a world that has moved.
It is still worth watching for a collapse. It is not worth reporting as progress, and it answers no question anybody in the operation is actually asking.
Count what still needs a person
The most useful single number is the share of cases that complete without human involvement, tracked weekly rather than quoted once at handover. It moves when the business moves, and its shape over time says more than its value on any given day.
- handled end to end
- escalated with a reason
- escalated with no reason given
The third category is the one to watch. It is where the system is failing without knowing that it is.
Time is the honest measure
The question a team actually has is whether the work is finished sooner, not whether the model was right about it.
Time from a case arriving to a case being resolved, measured across everything and not only the automated path, is the number that matches what people experience. An automation that is fast on its own cases and doubles the queue for the rest has not helped anybody.

Watch the reversals
Every decision a person overturns is a labelled error, delivered free. Counted by category, reversals are the clearest specification for the next change, and a category whose reversal rate climbs is the first sign the ground has shifted.
Confidence needs calibrating
A system that escalates below a threshold is only as good as that threshold, and the threshold was set before anyone had evidence. Comparing the confidence attached to a decision with whether it was later overturned is how you find out if the number means anything at all.
Cost per case, with people
Inference is rarely the expensive part. The honest figure includes the time spent on escalations, the time spent correcting decisions and the time spent maintaining the thing. Left out, it produces a saving that finance cannot find in the accounts.
Put it where the team looks
A dashboard nobody opens measures nothing. The numbers belong where the work already happens — a weekly figure in the channel the operations team reads, rather than a page that requires somebody to remember a URL.
Review it on a schedule
Once a month, with the person who owns the process, against the numbers from the month before. The purpose is not reporting. It is deciding what to change while the change is still small.
(nai™ — 11)
Insights & Research
Recent articles
Notes on automation, commerce and the software that has to carry both.

Aug 28, 2026
in /
Automation
Why operational automation stalls six months after launch

Aug 14, 2026
in /
Commerce
Moving a Shopify store to Hydrogen, and what actually changes

Aug 1, 2026
in /
AI systems
An assistant is only as good as what it is allowed to read

Jul 22, 2026
in /
Automation
The exceptions are the work, not the edge case

Jul 9, 2026
in /
Architecture
Nobody wants to pay for data work, and everybody pays for it

Jun 26, 2026
in /
Commerce
On a storefront, performance is not a technical concern

Jun 17, 2026
in /
Architecture
Build the field app for the worst signal it will ever see

Jun 5, 2026
in /
Strategy
When the spreadsheet is load-bearing, buying will not help
No hype. Just systems
Clarity beats automation
Decisions over demos
Designed for messy reality
Systems that hold under pressure

(nai™ — 15)
Work with us
Book a
free call
We build AI that changes
how your work runs — not how the demo looks.
We’ll map how your
workflows run, show
where AI pays off,
and leave you
with a clear plan.
No prep needed — we’ll steer the conversation
and keep it on what matters.




4 practices
7+ yrs
minimum, every engineer


