Measures, not vibes.
Calibrated and traceable to a standard. Uncertainty stated at every value.
A 16-year published research program
Metrology chapter. Barney & Barney (2024), Transdisciplinary Measurement through AI (De Gruyter), anchors the 16-year record and aligns with NIST AI 800-3. https://doi.org/10.1515/9783111036496-003 ยท Full record
Benchmarks score tasks. They do not measure performers.
the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.
See NIST AI 800-3, February 2026.
One horizontal ruler, 0 to 1.2 logits. Each bar is a published error at the decision point. 0.07 is the XLNC design target in validation. PILOT.
gold 0.10 good 0.25 Six Sigma 0.33 XLNC 0.07 design target (PILOT)
Measurand: conditional standard error at the decision or cut point, in logits (test-retest is the only fallback). Internal-consistency alpha, IRT marginal reliability, split-half, and composite reliability are not on this axis. Crosswalk: logit = 2 x (SE/SD); SD is about 1 MHC stage, about 2 logits (explicit modeling assumption). Metrology axis k=1.
Test names are text labels only, trademarks of their respective owners. XLNC is not affiliated with, endorsed by, or sponsored by any listed organization. Each value is that instrument's published conditional standard error or test-retest figure, crosswalked to this shared logit ruler. This is not a claim that XLNC is better than any named test.

The method is named by NIST AI 800-3 (February 2026) and published across a 16-year research program, including the Cambridge Handbook of Technology and Employee Behavior. Certification you can trust.