Raw benchmark scores cannot be averaged, subtracted, or compared across task sets. Rasch conjoint measurement yields a true interval scale, so the intervals between performers carry meaning.
the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.
NIST AI 800-3, February 2026. Quoted verbatim.
One frame of reference. Conditional standard error at the decision point, in logits. Metrology gold, good, and Six Sigma sit on the same axis as the familiar tests.
XLNC / AIM 0.07 design target (PILOT) 10:1 gold 0.10 4:1 good 0.25 Six Sigma 0.33 (Westgard)
Conditional SE at the decision point (logits, k=1). Tight on the left, wide on the right.
What this means: every mark is the same quantity, conditional standard error at the decision point, so a lower bar is a tighter measurement. XLNC 0.07 sits tighter than even the metrology gold 10:1 reference.
GMAT is held: no like-for-like published precision figure exists, so it is not plotted.
Physical instruments (vernier 0.02 mm, micrometer 0.01 mm, gauge blocks, SI metre) define minimum precision as a stated floor in their native units. The metrology discipline, a stated decision-quality uncertainty floor, is what the gold, good, and Six Sigma anchors carry onto this ruler. AIM is not a NIST or ISO standard; the claim is that the discipline is the same. Test names are trademarks of their owners; XLNC is not affiliated with, endorsed by, or sponsored by any listed organization. Crosswalk: logit = 2 x (SE/SD); within-sample SD is about 1 MHC stage, about 2 logits (explicit modeling assumption).
Certified 07-23 figures only. Axis measurand is conditional standard error at the decision or cut point, in logits, k=1. Internal-consistency alpha, IRT marginal reliability, split-half, and composite reliability are not on this ruler. SAT and WAIS-IV use the certified 07-23 test-retest fallback. GMAT is held because no like-for-like published precision figure exists. This is not a claim that XLNC is better than any named test.
| Instrument | Conditional SE (logits) | What is plotted |
|---|---|---|
| XLNC / AIM | 0.07 | 0.07 is the XLNC / AIM design target in validation. PILOT. |
| ACT | 0.34 | conditional SE at the cut |
| WAIS-IV | 0.40 | test-retest |
| SAT | 0.53 | test-retest |
| GRE Quant at 150 | 0.55 | conditional SE at the cut |
| GRE Verbal at 150 | 0.64 | conditional SE at the cut |
| USMLE Step 3 | 0.67 | global SEM on a known score SD |
| USMLE Step 2 CK | 0.80 | global SEM on a known score SD |
Interval-scale measurement, not raw scores.
One calibrated ruler measures every performer: human, AI agent, robot. One interval-scale metric. Stated uncertainty.
Raw benchmark scores rank. They do not measure. XLNC applies Rasch conjoint measurement to produce scores with equal intervals, so differences between performers mean the same thing everywhere on the scale.
Raw scores fail for a structural reason. A benchmark score is a count of correct answers, and counts are only valid as measurements when every item contributes equally. They never do. As the NIST AI 800-3 report states: "the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property." Change the item mix and the score changes, even when capability has not.
Rasch conjoint measurement solves this. It models item difficulty and performer ability on a single latent scale, converting ordinal response patterns into interval-scale estimates with stated standard errors. On the AIM ruler, we achieve a standard error of 0.07. CERTIFIED.
Construct maps make the scale interpretable. Our hallucination construct is a unidimensional Rasch construct. CERTIFIED. Our MHC stage-grain construct map locates developmental stages along the same logic of ordered, evenly spaced levels.
Three parallel forms are calibrated to one scale, so a performer measured on Form A is directly comparable to a performer measured on Form B. No re-anchoring. No re-norming.
Calibration is pre-registered. The calibration-validity program was registered on the Open Science Framework before data collection (project 3xf6r, registration j95ef), and the current measurement program is registered separately (fh8yd).
For the methodological foundation, see Cambridge Handbook Chapter 18: measuring minds without asking questions, the case for unobtrusive measurement. CERTIFIED.
Read the NIST Alignment
See the verbatim passage from NIST AI 800-3 and how interval-scale measurement resolves what raw scores cannot.
A 16-year published arc: inverted CAT (2010) to adaptive measurement (2016) to metrological psychometrics (2017) to AI + I-O (2019) to hybrid LLM metrology (2024) to AIM.



Fairness and reliability are not single numbers. The XLNC AIM engine reports judge severity and standard error per candidate level, because that is where score bias and noise actually hide.
When an LLM judge scores evaluation responses across ability or seniority bands, two failure modes can appear, and most reporting misses both. One is a severity slope: a judge that runs systematically harsher at one level than another. That slope is a bias signal. The other is a per-band SE blow-up: even when average severity looks flat across levels, the standard error inside one band can be inflated. Most vendors report one global reliability figure and one global fairness check. Both can pass while one level quietly fails.
Pilot dataset: Barney, Wind, & Krishna (2026). https://doi.org/10.21449/ijate.1788563
A judge can be fine mid-scale and blow up at a tail band. Global averages hide it. Per-band traces show it.
WITHOUT AIM the naive global |sev| looks like a near-tie (Qwen3 1.761 st, Llama 3.3 2.090 st, gap 0.329 st). The naive rule picks the wrong judge. WITH AIM the per-band trace is actionable: Llama 3.3 band 11 is 3.600 st (n=660, 2.880 logits), swing -2.402 st from band 8 to band 11, tail 1.72x its own average. The same judge that is fine mid-scale is the judge that fails at the tail band.
Shared legend (both panels): red = naive global rule, wrong judge, hidden tail. green = per-band trace, band 11 tail collapse exposed.
The calibration record is pre-registered and current.
The calibration-validity program is registered on the Open Science Framework (project 3xf6r, registration j95ef, public), with gates and verdicts frozen before collection. The current program, human and AI performance across four work domains on one ruler, is registered separately (fh8yd).
A 200-item wave of the reasoning-complexity bank was calibrated by four judges from different model families: item reliability 0.94, five to six distinct complexity strata, every judge above the agreement floor.
Stage boundaries are pinned to a certified reference calibrated on 747 human respondents (Dawson-Tunik, Commons, Wilson & Fischer, 2005). The current bank calibration reached EAP reliability 0.977 on 86 items, with stage ordering monotonic from stage 10 to 12.
The pre-registered unidimensionality gate failed on the 200-item wave, and the verdict was reported as a finding. A calibration that can catch its own defects is the point.