XLNC · Adaptive Intelligent Measurement

How AIM measures, without interrupting the work.

A judge’s score is not a measurement. AIM takes a sample you already have, probes it adaptively, removes each judge’s bias before it measures, and stops when it is certain. The result is a calibrated score with an honest interval.

Pilot tier, confirmatory calibration study in progress (OSF)

The unobtrusive adaptive workflow

No new test. No live exam. AIM measures from a sample the performer already produced, and adapts which prompt it asks the judge next, to learn the most per step.

1

Your sample input

A document, a transcript, a model output, text the performer already produced. Or pick a showcase dataset (e.g. MentalAlign-70k) to see it run. Nothing new to administer.

2

Adaptive prompt selection Fisher information

The inverted CAT chooses which prompt the judge scores next, the one that removes the most uncertainty about the performer’s stage. This is the adaptive part: it adapts the probe, not the performer.

3

Judges score, bias removed first MRCMLM / MFRM

Independent LLM judges score the sample. Each judge’s severity is estimated and subtracted before measurement, so the score reflects the performer, not which judge happened to read it.

4

It stops when it’s certain stopping rules

Measurement continues until a rule fires: 95-99% certainty of the MHC stage, a precision target (SE ≤ 0.30 logits), out of prompts, or a time / token budget. No wasted probes.

5

Calibrated θ + honest interval output

A bias-removed ability estimate on a logit ruler, with a confidence interval, never a bare number. Read on for how to interpret it as an MHC stage and a Vygotsky zone.

Bias is estimated, then removed, never ejected.

Every judge (human or LLM) scores with a stable severity bias, γ. AIM estimates it once from a calibration set and subtracts it before the CAT engine runs:

scorecorrected = scorerawγjudge   (in logit space, before adaptive measurement)

The judge is not thrown out, its bias is simply accounted for. θ stays clean regardless of which judge scored the sample.

What you read: precision for engineers, meaning for decisions

The same estimate is reported two ways, a logit value for precision, and an MHC stage with a Vygotsky zone for what to do next.

Raw, for the engineer

θ = +0.43 ± 0.22 logits
95% interval [−0.01, +0.87], uncertainty made visible, not hidden

Logits are the measurement layer: interval-scale, comparable across performers, judges, and time. This is what you log, diff, and build on.

MHC stage, for the decision

Stage 10 · Systematic

The performer coordinates two abstract systems and selects between them, they see the rules of one system as a special case of a higher-order frame.

For people: don’t assign Stage-12 work without support. For AI: Stage-11 prompting stretches it; Stage-12 breaks it.

The Vygotsky zone, what they’re ready to grow into

The most actionable part of the measurement. The Zone of Proximal Development is the band just above the current stage, the material a performer can reach with scaffolding, for a person or an AI.

Stage 8
Stage 9
Stage 10
current
Stage 11
ZPD
Stage 12

Ready for Stage 11 (Paradigmatic) with structured challenge and support. Aim coaching, training data, or prompt complexity at Stage 11, not Stage 12, and not below Stage 10. That is the difference between developing a performer and stalling one.