See a live measurement on the calibrated ruler: human and AI performers, one interval scale, uncertainty stated.
Watch AIM measure.
One calibrated ruler measures every performer: human, AI agent, robot. One interval-scale metric. Stated uncertainty.
A measurement session, start to finish, with nothing hidden.
Below is a live embed slot. When a session runs, you will see a performer measured on the same ruler we publish: item responses, scale positions, and a standard error attached to every score. On the AIM ruler that standard error is 0.07, and it is reported alongside each result rather than buried in a footnote.
Every published score carries two things: a stated standard error and traceable anchors on a unidimensional Rasch construct. That means any score can be checked, rechecked, and compared against any other score on the same instrument. Humans and AI agents occupy the same scale. A 0.42 is a 0.42 regardless of who or what produced it.
This is not a benchmark leaderboard. Raw benchmark scores do not have this property. As NIST AI 800-3 states verbatim: "the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property." Interval meaning is what makes comparison defensible, and defensibility is the point.
Request Access
Supporting line: Request access to schedule a measurement session and review full scoring documentation, including standard errors and anchor traceability, for your own performers.
Demo video slot reserved.
Footage of a live AIM measurement is being prepared. No simulated video is shown.