MEASUREMENT SCALE 2
How often a model asserts unsupported content, and where it stops.
When a model states something false with full confidence, you need to know how often that happens and where it stops, before a customer finds out. This scale measures a model's propensity to assert unsupported content, as a calibrated location on a cumulative construct, not a count of caught errors on one prompt set. It does not measure factual coverage or retrieval quality; a model can know less and hallucinate less, and this ruler will say so. A benchmark catches the hallucinations its items happened to touch. This ruler places the model on the construct with a stated standard error, so a new item form does not move the number. Calibrated; evidence accumulating.
Boundary. It does not measure factual coverage or retrieval quality; a model can know less and hallucinate less, and this ruler will say so.
Decision context. Feeds model selection and capability sufficiency decisions anywhere unsupported assertion is a liability.
Two-layer cumulative construct (Logic-2), Rasch/PCM family; in build on the certified harness. PILOT.
Measurement frame. Units are logits on the hallucination-propensity order; a one-logit difference is a constant odds ratio on asserting unsupported content. PILOT: the frame is calibrated, evidence accumulating.

Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.
| Band | Band label | Standard error | Certification implication |
|---|---|---|---|
| Bands 2+ | Two-layer cumulative scale, interior bands | SE 0.10 per band | PILOT: stated design precision, not yet certified. |
| Band 1 | Two-layer cumulative scale, first band | SE 0.15 | PILOT: stated design precision at the floor band, not yet certified. |
| -- | NO EVIDENCE Certified per-band calibration readout No certified per-band table exists yet for this scale; every band figure above is a PILOT design value, shown as such, and nothing uncertified is presented as certified. | -- | Shown and flagged, never removed. |
PILOT: calibrated, evidence accumulating.
Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.
Rows that fail the floor or lack evidence are shown and flagged, never removed.
The flagged row below is the honest-uncertainty state of this scale: the harness is in build, so the certified readout is absent and shown as such.
The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.
Status. In build on the certified harness. Not yet scoring externally. PILOT.
Scoring signature will mirror the MHC endpoint: submit a response set, receive a construct location with the standard error stated at the value.
POST /score {"scale": "hallucination", "responses": [...]} (PILOT; not yet open)
{"band": "<band>", "location_logit": "<stated>", "se": "<stated per band>"} (shape only; endpoint not yet live)
Nothing here is live. The shapes are stated so the contract is inspectable before the endpoint opens.
The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.