A judge’s score is not a measurement. AIM takes a sample you already have, probes it adaptively, removes each judge’s bias before it measures, and stops when it is certain. The result is a calibrated score with an honest interval.
Pilot tier, confirmatory calibration study in progress (OSF)No new test. No live exam. AIM measures from a sample the performer already produced, and adapts which prompt it asks the judge next, to learn the most per step.
A document, a transcript, a model output, text the performer already produced. Or pick a showcase dataset (e.g. MentalAlign-70k) to see it run. Nothing new to administer.
The inverted CAT chooses which prompt the judge scores next, the one that removes the most uncertainty about the performer’s stage. This is the adaptive part: it adapts the probe, not the performer.
Independent LLM judges score the sample. Each judge’s severity is estimated and subtracted before measurement, so the score reflects the performer, not which judge happened to read it.
Measurement continues until a rule fires: 95-99% certainty of the MHC stage, a precision target (SE ≤ 0.30 logits), out of prompts, or a time / token budget. No wasted probes.
A bias-removed ability estimate on a logit ruler, with a confidence interval, never a bare number. Read on for how to interpret it as an MHC stage and a Vygotsky zone.
Every judge (human or LLM) scores with a stable severity bias, γ. AIM estimates it once from a calibration set and subtracts it before the CAT engine runs:
The judge is not thrown out, its bias is simply accounted for. θ stays clean regardless of which judge scored the sample.
The same estimate is reported two ways, a logit value for precision, and an MHC stage with a Vygotsky zone for what to do next.
Logits are the measurement layer: interval-scale, comparable across performers, judges, and time. This is what you log, diff, and build on.
The performer coordinates two abstract systems and selects between them, they see the rules of one system as a special case of a higher-order frame.
For people: don’t assign Stage-12 work without support. For AI: Stage-11 prompting stretches it; Stage-12 breaks it.
The most actionable part of the measurement. The Zone of Proximal Development is the band just above the current stage, the material a performer can reach with scaffolding, for a person or an AI.
Ready for Stage 11 (Paradigmatic) with structured challenge and support. Aim coaching, training data, or prompt complexity at Stage 11, not Stage 12, and not below Stage 10. That is the difference between developing a performer and stalling one.