MEASUREMENT SCALE 6
Whether an agent finishes what it started, in order, within constraints.
When an agent runs a multi-step loop against your production systems, you need to know whether it completes the loop reliably, not whether it impressed someone on a demo set. This scale measures the reliability of agentic loops: does the agent finish what it started, in order, within the stated constraints. It does not measure the quality of any single step's output; a fast, fluent agent that drops step four is an unreliable agent, and this ruler says so. Provisional: decision-use prohibited for high-stakes calls until promotion criteria are met.
Boundary. It does not measure the quality of any single step's output; a fast, fluent agent that drops step four is an unreliable agent.
Decision context. Would feed agent-deployment assurance decisions; validated buyer pain H-CS-AGENTICLOOP-1.
Methodology start (canonical list T1-B). The category is crowded; the differentiator this instrument would carry is a stated SE on the reliability value. WATCH.
Measurement frame. No calibrated frame exists. Units and origin will be stated at calibration, on the same logit discipline as the other rulers.
Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.
No certified bands exist for this scale. A per-band SE table appears here only after a calibration wave is run and certified.
| Band | Band label | Standard error | Certification implication |
|---|---|---|---|
| -- | NO EVIDENCE Entire ruler No calibration wave has been run; there is no reliability location and no SE to state. Shown so the absence is on the record. | -- | Shown and flagged, never removed. |
WATCH: provisional; decision-use prohibited for high-stakes calls.
Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.
Rows that fail the floor or lack evidence are shown and flagged, never removed.
The flagged row below is the honest-uncertainty state of this scale: provisional, no certified measurement exists.
The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.
Status. None. WATCH scales expose no scoring endpoint.
No endpoint. Promotion to PILOT would publish the signature before any scoring opens.
(none; scale is provisional) (none; scale is provisional)
Nothing is asserted about endpoint behavior until the methodology build produces a calibrated instrument.
The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.