Cost per correctly verified outcome
currency/outcomeBetter when lowerEvery cost attributable to a run set — inference, tools, infrastructure and the human time spent approving, correcting or escalating — divided by the number of outcomes that passed an independent check.
How to measure it
Fix the task set and the verifier before running anything. Sum the cost of the whole set, failed attempts included. Divide by outcomes the verifier passed, not by runs that finished.
How to fake it
Narrow what counts as an outcome, or let the agent's own report stand as verification. Both shrink the denominator's honesty rather than the numerator's cost.
Read offACM-11