Error compounding kills multi-step pilots
Each step in a workflow is a separate bet. A naive 85% per-step success rate compounds to 14.2% over twelve steps - the shape of most real operational chains.
14.2% end-to-end at 12 steps
14.2%
Naive end-to-end
96.8%
Engineered end-to-end
Silent wrong actions are the real risk
On a 500-case AP exception queue, the naive pilot produced 40 silent wrong actions. Engineering converts failures into escalations and human gates - not hidden damage.
0 wrong actions on 500 cases
40
Naive wrong actions
0
Engineered wrong actions
Hardening beats swapping models
Pass rate on the same eval suite rises through engineering iterations - integrations, bounded scope, gates, regression cases - not a bigger frontier model.
54% → 92.3% on same model
54%
Baseline pass rate
92.3%
After hardening
Model upgrades need eval discipline
Over twelve months and three provider releases, unmonitored accuracy drifts from 92% to 71%. Eval-on-upgrade catches regressions before complaints do.
Floor holds above 90%
71%
Unmonitored floor
91.4%
Eval-on-upgrade floor



