XLNCXLNC Join the waitlist Talk to us

MEASUREMENT SCALE 6

Agent Eval / Agentic-Loop Reliability

WATCH

Whether an agent finishes what it started, in order, within constraints.

All measurement scales · One calibrated ruler · The method

S1. Construct definition

WATCH

When an agent runs a multi-step loop against your production systems, you need to know whether it completes the loop reliably, not whether it impressed someone on a demo set. This scale measures the reliability of agentic loops: does the agent finish what it started, in order, within the stated constraints. It does not measure the quality of any single step's output; a fast, fluent agent that drops step four is an unreliable agent, and this ruler says so. Provisional: decision-use prohibited for high-stakes calls until promotion criteria are met.

Boundary. It does not measure the quality of any single step's output; a fast, fluent agent that drops step four is an unreliable agent.

Decision context. Would feed agent-deployment assurance decisions; validated buyer pain H-CS-AGENTICLOOP-1.

S2. The Rasch/PCM ruler

Methodology start (canonical list T1-B). The category is crowded; the differentiator this instrument would carry is a stated SE on the reliability value. WATCH.

Measurement frame. No calibrated frame exists. Units and origin will be stated at calibration, on the same logit discipline as the other rulers.

S3. Stated SE per band

Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.

No certified bands exist for this scale. A per-band SE table appears here only after a calibration wave is run and certified.

BandBand labelStandard errorCertification implication
--NO EVIDENCE Entire ruler No calibration wave has been run; there is no reliability location and no SE to state. Shown so the absence is on the record.--Shown and flagged, never removed.

S4. Certification tier

WATCH

WATCH: provisional; decision-use prohibited for high-stakes calls.

Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.

S5. Honest-uncertainty display

Rows that fail the floor or lack evidence are shown and flagged, never removed.

The flagged row below is the honest-uncertainty state of this scale: provisional, no certified measurement exists.

The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.

S6. API and usage detail

Status. None. WATCH scales expose no scoring endpoint.

No endpoint. Promotion to PILOT would publish the signature before any scoring opens.

(none; scale is provisional)
(none; scale is provisional)

Nothing is asserted about endpoint behavior until the methodology build produces a calibrated instrument.

The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.

S7. Scientific references

  1. List SCALE_PRIORITIZATION_2026-08-31 rank 6: validated buyer pain H-CS-AGENTICLOOP-1; crowded eval-tooling category; displaced two rows by new candidates, not demoted on merit.