XLNC Join the waitlist

AI evals rank. They do not measure.

Raw benchmark scores cannot be averaged, subtracted, or compared across task sets. Rasch conjoint measurement yields a true interval scale, so the intervals between performers carry meaning.

the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.

NIST AI 800-3, February 2026. Quoted verbatim.

PILOT

One calibrated ruler

One frame of reference. Conditional standard error at the decision point, in logits. Metrology gold, good, and Six Sigma sit on the same axis as the familiar tests.

XLNC / AIM 0.07 design target (PILOT) 10:1 gold 0.10 4:1 good 0.25 Six Sigma 0.33 (Westgard)

Conditional SE at the decision point (logits, k=1). Tight on the left, wide on the right.

TighterWider
  • XLNC / AIM PILOTdesign target in validation0.07
  • 10:1 gold 0.10metrology decision-quality floor0.10
  • 4:1 good 0.25metrology decision-quality floor0.25
  • Six Sigma 0.33metrology decision-quality floor0.33
  • ACTconditional SE at the cut0.34
  • WAIS-IVtest-retest0.40
  • SATtest-retest0.53
  • GRE Quant at 150conditional SE at the cut0.55
  • GRE Verbal at 150conditional SE at the cut0.64
  • USMLE Step 3global SEM on a known score SD0.67
  • USMLE Step 2 CKglobal SEM on a known score SD0.80

What this means: every mark is the same quantity, conditional standard error at the decision point, so a lower bar is a tighter measurement. XLNC 0.07 sits tighter than even the metrology gold 10:1 reference.

GMAT is held: no like-for-like published precision figure exists, so it is not plotted.

Physical instruments (vernier 0.02 mm, micrometer 0.01 mm, gauge blocks, SI metre) define minimum precision as a stated floor in their native units. The metrology discipline, a stated decision-quality uncertainty floor, is what the gold, good, and Six Sigma anchors carry onto this ruler. AIM is not a NIST or ISO standard; the claim is that the discipline is the same. Test names are trademarks of their owners; XLNC is not affiliated with, endorsed by, or sponsored by any listed organization. Crosswalk: logit = 2 x (SE/SD); within-sample SD is about 1 MHC stage, about 2 logits (explicit modeling assumption).

Sources and method

Certified 07-23 figures only. Axis measurand is conditional standard error at the decision or cut point, in logits, k=1. Internal-consistency alpha, IRT marginal reliability, split-half, and composite reliability are not on this ruler. SAT and WAIS-IV use the certified 07-23 test-retest fallback. GMAT is held because no like-for-like published precision figure exists. This is not a claim that XLNC is better than any named test.

InstrumentConditional SE (logits)What is plotted
XLNC / AIM0.070.07 is the XLNC / AIM design target in validation. PILOT.
ACT0.34conditional SE at the cut
WAIS-IV0.40test-retest
SAT0.53test-retest
GRE Quant at 1500.55conditional SE at the cut
GRE Verbal at 1500.64conditional SE at the cut
USMLE Step 30.67global SEM on a known score SD
USMLE Step 2 CK0.80global SEM on a known score SD

Open the full ruler

Interval-scale measurement, not raw scores.

One calibrated ruler measures every performer: human, AI agent, robot. One interval-scale metric. Stated uncertainty.

Raw benchmark scores rank. They do not measure. XLNC applies Rasch conjoint measurement to produce scores with equal intervals, so differences between performers mean the same thing everywhere on the scale.

Raw scores fail for a structural reason. A benchmark score is a count of correct answers, and counts are only valid as measurements when every item contributes equally. They never do. As the NIST AI 800-3 report states: "the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property." Change the item mix and the score changes, even when capability has not.

Rasch conjoint measurement solves this. It models item difficulty and performer ability on a single latent scale, converting ordinal response patterns into interval-scale estimates with stated standard errors. On the AIM ruler, we achieve a standard error of 0.07. CERTIFIED.

Construct maps make the scale interpretable. Our hallucination construct is a unidimensional Rasch construct. CERTIFIED. Our MHC stage-grain construct map locates developmental stages along the same logic of ordered, evenly spaced levels.

Three parallel forms are calibrated to one scale, so a performer measured on Form A is directly comparable to a performer measured on Form B. No re-anchoring. No re-norming.

Calibration is pre-registered. The calibration-validity program was registered on the Open Science Framework before data collection (project 3xf6r, registration j95ef), and the current measurement program is registered separately (fh8yd).

For the methodological foundation, see Cambridge Handbook Chapter 18: measuring minds without asking questions, the case for unobtrusive measurement. CERTIFIED.

Read the NIST Alignment

See the verbatim passage from NIST AI 800-3 and how interval-scale measurement resolves what raw scores cannot.

The research program

A 16-year published arc: inverted CAT (2010) to adaptive measurement (2016) to metrological psychometrics (2017) to AI + I-O (2019) to hybrid LLM metrology (2024) to AIM.

  1. 2026 Barney, M., Wind, S., & Krishna, V. (2026). Using large language models to evaluate ethical persuasion text: A measurement modeling approach. IJATE, 13(1), 224-247. https://doi.org/10.21449/ijate.1788563
  2. 2024 Barney, M. & Barney, F. (2024). Transdisciplinary Measurement through AI: Hybrid metrology and psychometrics powered by large language models. In W.P. Fisher Jr. & L. Pendrill (Eds.), Models, Measurement, and Metrology Extending the Systeme International d'Unites. De Gruyter. https://doi.org/10.1515/9783111036496-003
  3. 2019 Barney, M.F. (2019). The Reciprocal Roles of Artificial Intelligence and Industrial-Organizational Psychology. In R.N. Landers (Ed.), Cambridge Handbook of Technology and Employee Behavior (pp. 3-21). Cambridge UP. https://doi.org/10.1017/9781108649636
  4. 2017 Barney, M.F. & Fisher, W. (2017). Avoiding AI Armageddon with Metrologically-Oriented Psychometrics. 18th International Congress of Metrology. https://doi.org/10.1051/metrology/201709005
  5. 2016 Barney, M.F. & Fisher, W.F. (2016). Adaptive Measurement and Assessment. Annual Review of Organizational Psychology and Organizational Behavior, 3, 469-490. https://doi.org/10.1146/annurev-orgpsych-041015-062329
  6. 2010 Barney, M.F. (2010). Inverted Computer-Adaptive Rasch Measurement: Prospects for Virtual and Actual Reality. IACAT, Arnhem. http://www.iacat.org/

Construct maps: what the ruler measures

Hallucination construct map, unidimensional Rasch
CERTIFIED Unidimensional Rasch construct for hallucination.
MHC reasoning complexity construct map
CERTIFIED MHC reasoning complexity construct map.

Three forms, one scale

Parallel forms on one interval scale
Parallel forms calibrated onto a single interval scale.

Judge severity by level

Fairness and reliability are not single numbers. The XLNC AIM engine reports judge severity and standard error per candidate level, because that is where score bias and noise actually hide.

When an LLM judge scores evaluation responses across ability or seniority bands, two failure modes can appear, and most reporting misses both. One is a severity slope: a judge that runs systematically harsher at one level than another. That slope is a bias signal. The other is a per-band SE blow-up: even when average severity looks flat across levels, the standard error inside one band can be inflated. Most vendors report one global reliability figure and one global fairness check. Both can pass while one level quietly fails.

Pilot dataset: Barney, Wind, & Krishna (2026). https://doi.org/10.21449/ijate.1788563

Different LLM judges for different capability levels.

A judge can be fine mid-scale and blow up at a tail band. Global averages hide it. Per-band traces show it.

WITHOUT AIM the naive global |sev| looks like a near-tie (Qwen3 1.761 st, Llama 3.3 2.090 st, gap 0.329 st). The naive rule picks the wrong judge. WITH AIM the per-band trace is actionable: Llama 3.3 band 11 is 3.600 st (n=660, 2.880 logits), swing -2.402 st from band 8 to band 11, tail 1.72x its own average. The same judge that is fine mid-scale is the judge that fails at the tail band.

Shared legend (both panels): red = naive global rule, wrong judge, hidden tail. green = per-band trace, band 11 tail collapse exposed.

WITHOUT AIM vs WITH AIM. Different LLM judges for different capability levels. PRELIMINARY PILOT. Two-condition contrast of judge severity on Reasoning Complexity (MHC Stage). WITHOUT AIM Qwen3 1.761 st, Llama 3.3 2.090 st, gap 0.329 st. Naive rule picks the wrong judge. WITH AIM Llama band 11 3.600 st n=660 2.880 logits, swing -2.402 st, tail 1.72x. Legend: red = naive global rule (wrong judge, hidden tail); green = per-band trace (band 11 tail collapse exposed). Sign reversal +0.28 to -3.60. WITHOUT AIM vs WITH AIM Different LLM judges for different capability levels. PRELIMINARY PILOT WITHOUT AIM One global number. The tail is hidden. Qwen3 1.761 st global |sev| Llama 3.3 2.090 st global |sev| gap 0.329 st a near-tie on the naive rule Pooled |sev| (bias spec 1.0 st) 1.0 Qwen3 1.761 Llama 3.3 2.090 Naive rule keeps Llama. That is the wrong judge. The tail stays hidden. WITH AIM Per-band trace. The tail band fails; the average hides it. band 11 3.600 st n=660 2.880 logits swing 8 to 11 -2.402 st tail 1.72x own average Llama 3.3 signed severity_hat -4.0 to +4.0 st 0 -4.0 +4.0 8 9 10 11 SE 0.1471 12 Tail collapse exposed. sign reversal +0.28 to -3.60. AIM picks a different judge. RED: naive global rule - wrong judge, hidden tail GREEN: per-band trace - band 11 tail collapse exposed Claim tier: PRELIMINARY PILOT. severity_hat = mean(judged stage minus target band), MHC stage units. Snapshot: 10,432-row judge_recred ledger, 2026-08-24. C1 swing is -2.402 st (band 8 to 11). AIM: built so every measure is traceable; anchoring in validation, arriving 2026.
PRELIMINARY PILOT Two-condition contrast. Signed severity_hat, MHC stage units. Snapshot: 10,432-row judge_recred ledger, 2026-08-24. Dataset: Barney, Wind, & Krishna (2026), doi:10.21449/ijate.1788563. Source: assets/img/with_aim_contrast_2cond.bbf3409d2f04.svg

Calibration benchmark

The calibration record is pre-registered and current.

CERTIFIED

Pre-registered on OSF

The calibration-validity program is registered on the Open Science Framework (project 3xf6r, registration j95ef, public), with gates and verdicts frozen before collection. The current program, human and AI performance across four work domains on one ruler, is registered separately (fh8yd).

CERTIFIED

Calibrated item bank

A 200-item wave of the reasoning-complexity bank was calibrated by four judges from different model families: item reliability 0.94, five to six distinct complexity strata, every judge above the agreement floor.

CERTIFIED

Anchored to human data

Stage boundaries are pinned to a certified reference calibrated on 747 human respondents (Dawson-Tunik, Commons, Wilson & Fischer, 2005). The current bank calibration reached EAP reliability 0.977 on 86 items, with stage ordering monotonic from stage 10 to 12.

CERTIFIED

A gate that can fail

The pre-registered unidimensionality gate failed on the 200-item wave, and the verdict was reported as a finding. A calibration that can catch its own defects is the point.