XLNCXLNC Join the waitlist Talk to us

MEASUREMENT SCALE 2

Hallucination Scale

PILOT

How often a model asserts unsupported content, and where it stops.

All measurement scales · One calibrated ruler · The method

S1. Construct definition

PILOT

When a model states something false with full confidence, you need to know how often that happens and where it stops, before a customer finds out. This scale measures a model's propensity to assert unsupported content, as a calibrated location on a cumulative construct, not a count of caught errors on one prompt set. It does not measure factual coverage or retrieval quality; a model can know less and hallucinate less, and this ruler will say so. A benchmark catches the hallucinations its items happened to touch. This ruler places the model on the construct with a stated standard error, so a new item form does not move the number. Calibrated; evidence accumulating.

Boundary. It does not measure factual coverage or retrieval quality; a model can know less and hallucinate less, and this ruler will say so.

Decision context. Feeds model selection and capability sufficiency decisions anywhere unsupported assertion is a liability.

S2. The Rasch/PCM ruler

Two-layer cumulative construct (Logic-2), Rasch/PCM family; in build on the certified harness. PILOT.

Measurement frame. Units are logits on the hallucination-propensity order; a one-logit difference is a constant odds ratio on asserting unsupported content. PILOT: the frame is calibrated, evidence accumulating.

Hallucination construct map card, Mari-Wilson hexagon, Logic-2 cumulative
CANDIDATE Hallucination (LLM Output Faithfulness Failure) - Logic-2 cumulative. Candidate construct map; public-reference pins gated on a certified referee.

S3. Stated SE per band

Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.

BandBand labelStandard errorCertification implication
Bands 2+Two-layer cumulative scale, interior bandsSE 0.10 per bandPILOT: stated design precision, not yet certified.
Band 1Two-layer cumulative scale, first bandSE 0.15PILOT: stated design precision at the floor band, not yet certified.
--NO EVIDENCE Certified per-band calibration readout No certified per-band table exists yet for this scale; every band figure above is a PILOT design value, shown as such, and nothing uncertified is presented as certified.--Shown and flagged, never removed.

S4. Certification tier

PILOT

PILOT: calibrated, evidence accumulating.

Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.

S5. Honest-uncertainty display

Rows that fail the floor or lack evidence are shown and flagged, never removed.

The flagged row below is the honest-uncertainty state of this scale: the harness is in build, so the certified readout is absent and shown as such.

The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.

S6. API and usage detail

Status. In build on the certified harness. Not yet scoring externally. PILOT.

Scoring signature will mirror the MHC endpoint: submit a response set, receive a construct location with the standard error stated at the value.

POST /score  {"scale": "hallucination", "responses": [...]}  (PILOT; not yet open)
{"band": "<band>", "location_logit": "<stated>", "se": "<stated per band>"}  (shape only; endpoint not yet live)

Nothing here is live. The shapes are stated so the contract is inspectable before the endpoint opens.

The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.

S7. Scientific references

  1. Plan HALLUCINATION_AIM_PIPELINE_PLAN_2026-08-17 (repo reports/plays): the build plan for the certified-harness instantiation of this scale.
  2. 2024 Barney, M. & Barney, F. (2024). Transdisciplinary Measurement through AI: Hybrid metrology and psychometrics powered by large language models. De Gruyter. https://doi.org/10.1515/9783111036496-003