MEASUREMENT SCALE 4
Certified evidence of what an agent got right in research outputs.
When an agent drafts the methods section, the results summary, or the lab record, you need certified evidence of what it got right, with the uncertainty stated, before a journal, a funder, or an auditor asks. This instrument measures the fidelity of agent-generated research outputs against certified reference procedures, on the text side reusing the adversarial-review and certification pipeline that already exists. It does not measure whether the underlying science is correct; it measures whether the agent's account of it holds up. Provisional: decision-use prohibited for high-stakes calls until promotion criteria are met.
Boundary. It does not measure whether the underlying science is correct; it measures whether the agent's account of it holds up.
Decision context. Would feed certification decisions for labs, journals, and funders facing agent-generated research outputs.
Not yet specified for calibration. The text side is designed to reuse the existing adversarial-review plus certification pipeline. WATCH: parked, promotion gated on CEO ratification, a scoping memo, a tested buyer hypothesis, and an explicit GO (rank-4 promotion rules).
Measurement frame. No calibrated frame exists. Units and origin will be stated at calibration, on the same logit discipline as the other rulers.
Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.
No certified bands exist for this scale. A per-band SE table appears here only after a calibration wave is run and certified.
| Band | Band label | Standard error | Certification implication |
|---|---|---|---|
| -- | NO EVIDENCE Entire ruler No calibration wave has been run; there is no bank, no band structure, and no SE to state. Shown so the absence is on the record, not implied away. | -- | Shown and flagged, never removed. |
WATCH: provisional; decision-use prohibited for high-stakes calls.
Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.
Rows that fail the floor or lack evidence are shown and flagged, never removed.
The flagged row below is the honest-uncertainty state of this scale: parked, no measurement exists.
The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.
Status. None. WATCH scales expose no scoring endpoint.
No endpoint. Promotion to PILOT would publish the signature before any scoring opens.
(none; scale is parked) (none; scale is parked)
Per the promotion rules, no build, no spend, and no outreach occur while this scale is parked.
The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.