XLNCXLNC Try the Demo
CERTIFIED

Measurement reduces uncertainty for better decisions.

Measures, not vibes.

Calibrated and traceable to a standard. Uncertainty stated at every value.

Measurement narrows uncertainty against a standard A ruler draws. A wide credibility interval overshoots a standard band, then narrows and lands inside it. The Plumb Line mark completes and holds. STANDARD
Measurement reduces uncertainty for better decisions.

A 16-year published research program

Full citations

Metrology chapter. Barney & Barney (2024), Transdisciplinary Measurement through AI (De Gruyter), anchors the 16-year record and aligns with NIST AI 800-3. https://doi.org/10.1515/9783111036496-003 ยท Full record

Benchmarks score. AIM measures.

Benchmarks score tasks. They do not measure performers.

  • Benchmarks score tasks.
  • They do not measure performers.
  • A 0.85 today is not a 0.85 tomorrow.
  • No two systems share a scale.
  • One unidimensional Rasch ruler.
  • Humans, AI agents, robots.
  • Same interval-scale metric.
  • Measurement error stated at every value.
One ruler across human, AI agent, robot
Humans and AI agents sit on the same interval scale. Uncertainty is stated, never hidden.

Certified precision, published method.

  • Standard error 0.07 on the AIM ruler. CERTIFIED.
  • 0.07 is the design target in validation. PILOT.
  • Method published: Chapter 18, Cambridge Handbook.
  • NIST AI 800-3: intervals between LLM capabilities have meaning independent of the tasks selected.
  • Raw benchmark scores do not.
the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.

See NIST AI 800-3, February 2026.

PILOT

One calibrated ruler. Every test's error bar in one view.

One horizontal ruler, 0 to 1.2 logits. Each bar is a published error at the decision point. 0.07 is the XLNC design target in validation. PILOT.

gold 0.10 good 0.25 Six Sigma 0.33 XLNC 0.07 design target (PILOT)

XLNC PILOTdesign target in validation
0.07
ACTconditional SE at the cut
0.34
Wonderlictest-retest
0.49
SAT Totaltest-retest
0.53
GRE Quant at 150conditional SE at the cut
0.55
NEO-PI-Rtest-retest
0.57
GRE Verbal at 150conditional SE at the cut
0.64
USMLE Step 3global SEM on a known score SD
0.67
MBTItest-retest
0.69
USMLE Step 2 CKglobal SEM on a known score SD
0.80
CliftonStrengthstest-retest
1.10

Measurand: conditional standard error at the decision or cut point, in logits (test-retest is the only fallback). Internal-consistency alpha, IRT marginal reliability, split-half, and composite reliability are not on this axis. Crosswalk: logit = 2 x (SE/SD); SD is about 1 MHC stage, about 2 logits (explicit modeling assumption). Metrology axis k=1.

Full table and method note

Measure free. Certify paid.

  • Measurement is free.
  • Certification is paid.
  • Screen, coach, or benchmark at no cost.
  • Audit, procurement, or litigation: paid tier.
  • You pay only when the number has to survive scrutiny.
Standard-error ruler from 0.30 screening to 0.07 ceiling
Free measures sit at the wide end. Certified precision sits at 0.07.

Authority you can check

The method is named by NIST AI 800-3 (February 2026) and published across a 16-year research program, including the Cambridge Handbook of Technology and Employee Behavior. Certification you can trust.