One fixed maturity scale
A score changes only when the evidence crosses a threshold.
The ideal target is level 4 across the operational dimensions and level 5 for human outcome. Autonomy is a deployment mode, not another maturity level.
- 0
Not demonstrated
No credible public evidence that the system can perform the job.
- 1
Demonstrable
Selected examples show it can happen; reliability is unknown.
- 2
Repeatable
Held-out bounded tasks pass under declared conditions.
- 3
Usable
Intended practitioners can steer, inspect, correct, and accept it.
- 4
Deliverable
It survives production, access, handoff, and later update.
- 5
Outcome-proven
It improves the intended human outcome against the right baseline.
No aggregate score.
A mean would let strong code generation conceal weak integrity or reader outcomes. Confidence is reported separately from maturity.