Raw benchmark scores cannot be averaged, subtracted, or compared across task sets. Rasch conjoint measurement yields a true interval scale, so the intervals between performers carry meaning.
the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.
NIST AI 800-3, February 2026. Quoted verbatim.
Every plotted bar is that instrument's own published number, converted to one shared logit ruler of measurement uncertainty. 0.07 is the XLNC design target in validation. PILOT.
gold 0.10 good 0.25 Six Sigma 0.33 XLNC 0.07 design target (PILOT)
| Instrument | Logits | What is plotted |
|---|---|---|
| XLNC | 0.07 | 0.07 is the XLNC design target in validation. PILOT. |
| ACT | 0.34 | conditional SE at the cut |
| Wonderlic | 0.49 | test-retest |
| SAT Total | 0.53 | test-retest |
| GRE Quant at 150 | 0.55 | conditional SE at the cut |
| NEO-PI-R | 0.57 | test-retest |
| GRE Verbal at 150 | 0.64 | conditional SE at the cut |
| USMLE Step 3 | 0.67 | global SEM on a known score SD |
| MBTI | 0.69 | test-retest |
| USMLE Step 2 CK | 0.80 | global SEM on a known score SD |
| CliftonStrengths | 1.10 | test-retest |
Measurand: conditional standard error at the decision or cut point, in logits (test-retest is the only fallback). Internal-consistency alpha, IRT marginal reliability, split-half, and composite reliability are not on this axis. Crosswalk: logit = 2 x (SE/SD); SD is about 1 MHC stage, about 2 logits (explicit modeling assumption). Metrology axis k=1.
Test names are text labels only, trademarks of their respective owners. XLNC is not affiliated with, endorsed by, or sponsored by any listed organization. Each value is that instrument's published conditional standard error or test-retest figure, crosswalked to this shared logit ruler. This is not a claim that XLNC is better than any named test.
Interval-scale measurement, not raw scores.
One calibrated ruler measures every performer: human, AI agent, robot. One interval-scale metric. Stated uncertainty.
Raw benchmark scores rank. They do not measure. XLNC applies Rasch conjoint measurement to produce scores with equal intervals, so differences between performers mean the same thing everywhere on the scale.
Raw scores fail for a structural reason. A benchmark score is a count of correct answers, and counts are only valid as measurements when every item contributes equally. They never do. As the NIST AI 800-3 report states: "the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property." Change the item mix and the score changes, even when capability has not.
Rasch conjoint measurement solves this. It models item difficulty and performer ability on a single latent scale, converting ordinal response patterns into interval-scale estimates with stated standard errors. On the AIM ruler, we achieve a standard error of 0.07. CERTIFIED.
Construct maps make the scale interpretable. Our hallucination construct is a unidimensional Rasch construct. CERTIFIED. Our MHC stage-grain construct map locates developmental stages along the same logic of ordered, evenly spaced levels.
Three parallel forms are calibrated to one scale, so a performer measured on Form A is directly comparable to a performer measured on Form B. No re-anchoring. No re-norming.
The Coaching Supervision stack serves as our calibration benchmark. It is Rasch-analyzed, not AIM-anchored. RASCH-ANALYZED, NOT AIM-ANCHORED.
For the methodological foundation, see Cambridge Handbook Chapter 18: measuring minds without asking questions, the case for unobtrusive measurement. CERTIFIED.
Read the NIST Alignment
See the verbatim passage from NIST AI 800-3 and how interval-scale measurement resolves what raw scores cannot.
A 16-year published arc: inverted CAT (2010) to adaptive measurement (2016) to metrological psychometrics (2017) to AI + I-O (2019) to hybrid LLM metrology (2024) to AIM.



Fairness and reliability are not single numbers. The XLNC AIM engine reports judge severity and standard error per candidate level, because that is where score bias and noise actually hide.
When an LLM judge scores evaluation responses across ability or seniority bands, two failure modes can appear, and most reporting misses both. One is a severity slope: a judge that runs systematically harsher at one level than another. That slope is a bias signal. The other is a per-band SE blow-up: even when average severity looks flat across levels, the standard error inside one band can be inflated. Most vendors report one global reliability figure and one global fairness check. Both can pass while one level quietly fails.
Pilot dataset: Barney, Wind, & Krishna (2026). https://doi.org/10.21449/ijate.1788563
The judge your averages would pick. The judge your tails need.
WITHOUT AIM the naive global |sev| looks like a near-tie (Qwen3 1.761 st, Llama 3.3 2.090 st, gap 0.329 st). The naive rule picks the wrong judge. WITH AIM the per-band trace is actionable: Llama 3.3 band 11 is 3.600 st (n=660, 2.880 logits), swing -2.402 st from band 8 to band 11, tail 1.72x its own average.
SME-validated and Rasch-analyzed as a finished human instrument. Discipline label: RASCH-ANALYZED, NOT AIM-ANCHORED.