One fixed maturity scale

A score changes only when the evidence crosses a threshold.

The ideal target is level 4 across the operational dimensions and level 5 for human outcome. Autonomy is a deployment mode, not another maturity level.

  1. 0

    Not demonstrated

    No credible public evidence that the system can perform the job.

  2. 1

    Demonstrable

    Selected examples show it can happen; reliability is unknown.

  3. 2

    Repeatable

    Held-out bounded tasks pass under declared conditions.

  4. 3

    Usable

    Intended practitioners can steer, inspect, correct, and accept it.

  5. 4

    Deliverable

    It survives production, access, handoff, and later update.

  6. 5

    Outcome-proven

    It improves the intended human outcome against the right baseline.

No aggregate score.

A mean would let strong code generation conceal weak integrity or reader outcomes. Confidence is reported separately from maturity.