These are real, frozen measurement sessions produced by the AIM measurement and evaluation engine against the production item bank: every item selection, every theta update, and every standard error below is engine output, captured turn by turn. They are recorded sessions, not a live measurement. Live, interactive measurement on your own performers ships in Phase 2.
Each session is a computerized adaptive measurement run: the engine picks the next item from the calibrated bank based on the current estimate, scores the response, updates the capability estimate (theta) and its standard error, and stops when the stopping rules are met. Demo sessions run at 90% confidence for speed; production certification uses stricter stopping. The chart shows theta narrowing as evidence accumulates; the band is plus or minus one standard error. No number on this page is published without its uncertainty.
Quality first. Cost next. Both on one ruler. For each MHC band, the row shows the cheapest credentialed model whose measured capability ceiling reaches the band and whose per-band measurement precision meets the frozen quality floor (SE ≤ 0.20). Cost per measure and mean latency come from the same probe telemetry; no number is published without its uncertainty. The floor is not user-degradable: the picker chooses among acceptable panels only.
| MHC band | Cheapest model clearing the floor (ceiling, band SE, evidence) | Cost per measure | Mean latency |
|---|---|---|---|
| S8 | deepseek-v4-flash-0731ceiling S11; band severity SE 0.104 (n=270); offset -1.19 MHC; measured 2026-08-19 · also clears: z-ai-glm-5-3 ($0.0016), grok-4-6 ($0.0064) | $0.0004 | 26.8 s |
| S9 | deepseek-v4-flash-0731ceiling S11; band severity SE 0.127 (n=270); offset -0.42 MHC; measured 2026-08-19 · also clears: z-ai-glm-5-3 ($0.0016), grok-4-6 ($0.0064) | $0.0004 | 26.8 s |
| S10 | deepseek-v4-flash-0731ceiling S11; band severity SE 0.133 (n=266); offset -1.57 MHC; measured 2026-08-19 · also clears: z-ai-glm-5-3 ($0.0016), grok-4-6 ($0.0064) | $0.0004 | 26.8 s |
| S11 | deepseek-v4-flash-0731ceiling S11; band severity SE 0.108 (n=265); offset -2.74 MHC; measured 2026-08-19 · also clears: z-ai-glm-5-3 ($0.0016), grok-4-6 ($0.0064) | $0.0004 | 26.8 s |
| S12 | grok-4-6ceiling S14; band severity SE 0.019 (n=225); offset -1.03 MHC; measured 2026-08-22 | $0.0064 | 20.2 s |
Ceilings: RCC-v1.3 generation probe under the monotonicity rule (FRESH registry rows only). Severity + SE: per-band judge evidence from the credential administrations of 2026-08-19 and 2026-08-22. Cost + latency: per-call telemetry from the same ledgers. This is a readout over existing measurement, not a new estimate.
Recorded measurement session. 172 items administered by the adaptive engine; stop rule: Precision + discipline coverage satisfied (SE=0.139, 21 disciplines covered, 3 anchor stages).
Replay of the recorded session: the playhead walks the 172 administered items in order; the theta line, the plus or minus one SE band, and the readout show the estimate at each turn.
Artifact canned-cat-mid-2026-08-27 · bank aim/data/prompt_bank.db (4212 items, sha256 6e91c0f758061114) · engine aim/cat_session.py (CATSession + AdaptiveNavigator + Bayesian estimator + StoppingRules) · generated 2026-08-27.
Recorded measurement session. 139 items administered by the adaptive engine; stop rule: Stage hypothesis confirmed: 12 (p=0.993).
Replay of the recorded session: the playhead walks the 139 administered items in order; the theta line, the plus or minus one SE band, and the readout show the estimate at each turn.
Artifact canned-cat-high-2026-08-27 · bank aim/data/prompt_bank.db (4212 items, sha256 af70aa616e4a7b66) · engine aim/cat_session.py (CATSession + AdaptiveNavigator + Bayesian estimator + StoppingRules) · generated 2026-08-27.
These sessions are recorded. The live path (same engine, real time, your evidence) is in Phase 2 and is CEO-scheduled. Book a live demo to see a live session.