XLNCXLNC

The published record behind every number on this surface. Scale-level citations live on the per-scale pages; this page carries the program-level record.

NIST AI 800-3#

NIST AI 800-3, Expanding the AI Evaluation Toolbox with Statistical Models (February 2026), names the interval-scale property AIM is built on. Verbatim:

the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.

The requirement-by-requirement scorecard is on the NIST Alignment page of the main site.

Published method#

The measurement method is published, not proprietary black box:

  • Barney, M.F. (2019). The Reciprocal Roles of Artificial Intelligence and Industrial-Organizational Psychology. Chapter 18 in R.N. Landers (Ed.), Cambridge Handbook of Technology and Employee Behavior (pp. 3-21). Cambridge University Press. doi.org/10.1017/9781108649636
  • Barney, M. and Barney, F. (2024). Transdisciplinary Measurement through AI: Hybrid metrology and psychometrics powered by large language models. In W.P. Fisher Jr. and L. Pendrill (Eds.), Models, Measurement, and Metrology Extending the Systeme International d'Unites. De Gruyter. doi.org/10.1515/9783111036496-003
  • Barney, M.F. and Fisher, W. (2017). Avoiding AI Armageddon with Metrologically-Oriented Psychometrics. 18th International Congress of Metrology. doi.org/10.1051/metrology/201709005
  • Barney, M.F. and Fisher, W.F. (2016). Adaptive Measurement and Assessment. Annual Review of Organizational Psychology and Organizational Behavior, 3, 469-490. doi.org/10.1146/annurev-orgpsych-041015-062329
  • Barney, M., Wind, S., and Krishna, V. (2026). Using large language models to evaluate ethical persuasion text: A measurement modeling approach. IJATE, 13(1), 224-247. doi.org/10.21449/ijate.1788563
  • Barney, M.F. (2010). Inverted Computer-Adaptive Rasch Measurement: Prospects for Virtual and Actual Reality. IACAT, Arnhem. iacat.org

OSF registrations#

The calibration program is pre-registered on the Open Science Framework, with gates and verdicts frozen before data collection:

  • Calibration-validity program: OSF project 3xf6r, registration j95ef (public). osf.io/j95ef
  • Current program, human and AI performance across four work domains on one ruler: registration fh8yd. osf.io/fh8yd

Traceability framework#

Construct specification follows the Mari and Wilson metrology-psychometric hexagon framework: every construct specified from measurand to public value. A co-authored paper on MASEMS (Metasystematic Assessment of Systematic Evaluations and Meta-Syntheses; Thornton, Barney and Fisher) is in preparation.

How to cite a number from AIM#

Cite the band location, the standard error stated at that location, the scale name, the certification tier, and the calibration vintage. A number without its SE and vintage is not an AIM number and should not be attributed to AIM.