XLNCXLNC Try the Demo

AI evals rank. They do not measure.

Raw benchmark scores cannot be averaged, subtracted, or compared across task sets. Rasch conjoint measurement yields a true interval scale, so the intervals between performers carry meaning.

the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.

NIST AI 800-3, February 2026. Quoted verbatim.

PILOT

One calibrated ruler. Every test's error bar in one view.

Every plotted bar is that instrument's own published number, converted to one shared logit ruler of measurement uncertainty. 0.07 is the XLNC design target in validation. PILOT.

gold 0.10 good 0.25 Six Sigma 0.33 XLNC 0.07 design target (PILOT)

XLNC PILOTdesign target in validation
0.07
ACTconditional SE at the cut
0.34
Wonderlictest-retest
0.49
SAT Totaltest-retest
0.53
GRE Quant at 150conditional SE at the cut
0.55
NEO-PI-Rtest-retest
0.57
GRE Verbal at 150conditional SE at the cut
0.64
USMLE Step 3global SEM on a known score SD
0.67
MBTItest-retest
0.69
USMLE Step 2 CKglobal SEM on a known score SD
0.80
CliftonStrengthstest-retest
1.10
InstrumentLogitsWhat is plotted
XLNC0.070.07 is the XLNC design target in validation. PILOT.
ACT0.34conditional SE at the cut
Wonderlic0.49test-retest
SAT Total0.53test-retest
GRE Quant at 1500.55conditional SE at the cut
NEO-PI-R0.57test-retest
GRE Verbal at 1500.64conditional SE at the cut
USMLE Step 30.67global SEM on a known score SD
MBTI0.69test-retest
USMLE Step 2 CK0.80global SEM on a known score SD
CliftonStrengths1.10test-retest

Measurand: conditional standard error at the decision or cut point, in logits (test-retest is the only fallback). Internal-consistency alpha, IRT marginal reliability, split-half, and composite reliability are not on this axis. Crosswalk: logit = 2 x (SE/SD); SD is about 1 MHC stage, about 2 logits (explicit modeling assumption). Metrology axis k=1.

Interval-scale measurement, not raw scores.

One calibrated ruler measures every performer: human, AI agent, robot. One interval-scale metric. Stated uncertainty.

Raw benchmark scores rank. They do not measure. XLNC applies Rasch conjoint measurement to produce scores with equal intervals, so differences between performers mean the same thing everywhere on the scale.

Raw scores fail for a structural reason. A benchmark score is a count of correct answers, and counts are only valid as measurements when every item contributes equally. They never do. As the NIST AI 800-3 report states: "the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property." Change the item mix and the score changes, even when capability has not.

Rasch conjoint measurement solves this. It models item difficulty and performer ability on a single latent scale, converting ordinal response patterns into interval-scale estimates with stated standard errors. On the AIM ruler, we achieve a standard error of 0.07. CERTIFIED.

Construct maps make the scale interpretable. Our hallucination construct is a unidimensional Rasch construct. CERTIFIED. Our MHC stage-grain construct map locates developmental stages along the same logic of ordered, evenly spaced levels.

Three parallel forms are calibrated to one scale, so a performer measured on Form A is directly comparable to a performer measured on Form B. No re-anchoring. No re-norming.

The Coaching Supervision stack serves as our calibration benchmark. It is Rasch-analyzed, not AIM-anchored. RASCH-ANALYZED, NOT AIM-ANCHORED.

For the methodological foundation, see Cambridge Handbook Chapter 18: measuring minds without asking questions, the case for unobtrusive measurement. CERTIFIED.

Read the NIST Alignment

See the verbatim passage from NIST AI 800-3 and how interval-scale measurement resolves what raw scores cannot.

The research program

A 16-year published arc: inverted CAT (2010) to adaptive measurement (2016) to metrological psychometrics (2017) to AI + I-O (2019) to hybrid LLM metrology (2024) to AIM.

  1. 2026 Barney, M., Wind, S., & Krishna, V. (2026). Using large language models to evaluate ethical persuasion text: A measurement modeling approach. IJATE, 13(1), 224-247. https://doi.org/10.21449/ijate.1788563
  2. 2024 Barney, M. & Barney, F. (2024). Transdisciplinary Measurement through AI: Hybrid metrology and psychometrics powered by large language models. In W.P. Fisher Jr. & L. Pendrill (Eds.), Models, Measurement, and Metrology Extending the Systeme International d'Unites. De Gruyter. https://doi.org/10.1515/9783111036496-003
  3. 2019 Barney, M.F. (2019). The Reciprocal Roles of Artificial Intelligence and Industrial-Organizational Psychology. In R.N. Landers (Ed.), Cambridge Handbook of Technology and Employee Behavior (pp. 3-21). Cambridge UP. https://doi.org/10.1017/9781108649636
  4. 2017 Barney, M.F. & Fisher, W. (2017). Avoiding AI Armageddon with Metrologically-Oriented Psychometrics. 18th International Congress of Metrology. https://doi.org/10.1051/metrology/201709005
  5. 2016 Barney, M.F. & Fisher, W.F. (2016). Adaptive Measurement and Assessment. Annual Review of Organizational Psychology and Organizational Behavior, 3, 469-490. https://doi.org/10.1146/annurev-orgpsych-041015-062329
  6. 2010 Barney, M.F. (2010). Inverted Computer-Adaptive Rasch Measurement: Prospects for Virtual and Actual Reality. IACAT, Arnhem. http://www.iacat.org/

Construct maps: what the ruler measures

Hallucination hexagon map, unidimensional Rasch
CERTIFIED Unidimensional Rasch construct for hallucination.
MHC reasoning complexity hexagon map
CERTIFIED MHC reasoning complexity hexagon map.

Three forms, one scale

Parallel forms on one interval scale
Parallel forms calibrated onto a single interval scale.

Judge severity by level

Fairness and reliability are not single numbers. The XLNC AIM engine reports judge severity and standard error per candidate level, because that is where score bias and noise actually hide.

When an LLM judge scores evaluation responses across ability or seniority bands, two failure modes can appear, and most reporting misses both. One is a severity slope: a judge that runs systematically harsher at one level than another. That slope is a bias signal. The other is a per-band SE blow-up: even when average severity looks flat across levels, the standard error inside one band can be inflated. Most vendors report one global reliability figure and one global fairness check. Both can pass while one level quietly fails.

Pilot dataset: Barney, Wind, & Krishna (2026). https://doi.org/10.21449/ijate.1788563

WITHOUT AIM vs WITH AIM

The judge your averages would pick. The judge your tails need.

WITHOUT AIM the naive global |sev| looks like a near-tie (Qwen3 1.761 st, Llama 3.3 2.090 st, gap 0.329 st). The naive rule picks the wrong judge. WITH AIM the per-band trace is actionable: Llama 3.3 band 11 is 3.600 st (n=660, 2.880 logits), swing -2.402 st from band 8 to band 11, tail 1.72x its own average.

WITHOUT AIM vs WITH AIM. The judge your averages would pick. The judge your tails need. PRELIMINARY PILOT. Two-condition contrast of judge severity on Reasoning Complexity (MHC Stage). WITHOUT AIM Qwen3 1.761 st, Llama 3.3 2.090 st, gap 0.329 st. Naive rule picks the wrong judge. WITH AIM Llama band 11 3.600 st n=660 2.880 logits, swing -2.402 st, tail 1.72x. Sign reversal +0.28 to -3.60. WITHOUT AIM vs WITH AIM The judge your averages would pick. The judge your tails need. PRELIMINARY PILOT WITHOUT AIM Naive global |sev|. Looks like a near-tie. Qwen3 1.761 st global |sev| Llama 3.3 2.090 st global |sev| gap 0.329 st a near-tie on the naive rule Pooled |sev| (bias spec 1.0 st) 1.0 Qwen3 1.761 Llama 3.3 2.090 Naive rule keeps Llama. That is the wrong judge. The tail stays hidden. WITH AIM Per-band control. Llama tail collapse is exposed. band 11 3.600 st n=660 2.880 logits swing 8 to 11 -2.402 st tail 1.72x own average Llama 3.3 signed severity_hat -4.0 to +4.0 st 0 -4.0 +4.0 8 9 10 11 SE 0.1471 12 Tail collapse exposed. sign reversal +0.28 to -3.60. AIM picks a different judge. Claim tier: PRELIMINARY PILOT. severity_hat = mean(judged stage minus target band), MHC stage units. Snapshot: 10,432-row judge_recred ledger, 2026-08-24. C1 swing is -2.402 st (band 8 to 11). AIM: built so every measure is traceable; anchoring in validation, arriving 2026.
PRELIMINARY PILOT Two-condition contrast. Signed severity_hat, MHC stage units. Snapshot: 10,432-row judge_recred ledger, 2026-08-24. Dataset: Barney, Wind, & Krishna (2026), doi:10.21449/ijate.1788563. Source: assets/img/with_aim_contrast_2cond.41c14debf4f5.svg

Calibration benchmark

CERTIFIED

Coaching Supervision stack

SME-validated and Rasch-analyzed as a finished human instrument. Discipline label: RASCH-ANALYZED, NOT AIM-ANCHORED.