XLNCXLNC Try the Demo

NIST names the method. AIM implements it at measurement grade.

Federal guidance on AI evaluation now names the interval-scale property that raw benchmark scores lack.

the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.

NIST AI 800-3, February 2026. Quoted verbatim. Source: nist.gov publications.

NIST names the method. AIM implements it at measurement grade.

One calibrated ruler measures every performer: human, AI agent, robot. One interval-scale metric. Stated uncertainty.

The National Institute of Standards and Technology describes what benchmark scores cannot do, and what a latent scale can. The distinction matters for anyone procuring AI systems on evidence rather than vendor claims. Source: NIST AI 800-3.

"the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property."

Raw benchmark scores are task-bound. Change the task mix and the number changes, with no way to say by how much or what it means. An interval scale removes that dependency. Differences between performers carry meaning in their own right, so a gap of one unit means the same thing at the top of the scale as at the bottom.

For buyers, this converts comparison into measurement. The AIM ruler reports SE 0.07 (CERTIFIED), built on a unidimensional Rasch hallucination construct (CERTIFIED), with methodology published as Chapter 18 of the Cambridge Handbook (CERTIFIED). Every score ships with stated uncertainty, so procurement teams can set thresholds, compare across vendors and versions, and defend the decision later.

Proposed tiers reflect this discipline (all PROPOSED): SE 0.30 screening and coaching at $150 to 400; SE 0.20 default analysis at $1,500 per analysis; SE 0.14 defensible at $3,000 to 5,000; SE 0.10 certification at $250,000 to 750,000 per year. Tighter uncertainty, higher assurance, priced accordingly.

Read the Science

See the construct, the calibration, and the published chapter behind every score.

What this means for buyers

On an interval scale, a two-point gap means the same amount everywhere on the ruler, across any task set. Raw scores cannot promise that. AIM publishes interval-scale scores with stated standard errors and traceable anchors.