MEASUREMENT SCALE 5
How much of an agent's stated rules survive repeated context compaction.
When your agent has been running for weeks and its context has been compacted a dozen times, you need to know how much of its safety rules are still intact, not whether the demo worked in March. This instrument measures the fraction of stated operating rules an agent retains across repeated context compactions, as a coefficient with a stated standard error. It does not measure general capability drift or model version change; it measures rule survival under compression. Provisional: decision-use prohibited for high-stakes calls until the compaction harness is built and the retention baseline is measured.
Boundary. It does not measure general capability drift or model version change; it measures rule survival under compression.
Decision context. Would feed fleet-assurance decisions for enterprise platform teams running long-running agents.
Not yet built. The instrument requires a repeatable N-round compaction harness with a calibrated retention panel before any certified measure exists (rank-5 promotion rules). WATCH.
Measurement frame. No calibrated frame exists. The coefficient and its standard error will be stated at calibration.
Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.
No certified bands exist for this scale. The illustrative figure in the canonical source (rules 10% intact after 5 compactions) is NOT a verified measurement and is excluded from this page per claim-tier discipline.
| Band | Band label | Standard error | Certification implication |
|---|---|---|---|
| -- | NO EVIDENCE Entire ruler No compaction harness exists and no retention baseline is measured; there is no coefficient and no SE to state. Shown so the absence is on the record. | -- | Shown and flagged, never removed. |
WATCH: provisional; decision-use prohibited for high-stakes calls.
Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.
Rows that fail the floor or lack evidence are shown and flagged, never removed.
The flagged row below is the honest-uncertainty state of this scale: parked, no measurement exists.
The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.
Status. None. WATCH scales expose no scoring endpoint.
No endpoint. Promotion to PILOT would publish the signature before any scoring opens.
(none; scale is parked) (none; scale is parked)
Per the promotion rules, no build, no spend, and no outreach occur while this scale is parked.
The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.