Capability tells you what a model can reach. Severity tells you what it will honestly say about what it sees. Neither alone is enough, and the interaction between them is where the routing mistakes hide.
2026-08-28 · measurement and evaluation
Your engineering team runs the same evaluation on three frontier models. All three ace the hardest reasoning tasks you throw at them. All three score within a point of each other on the leaderboard. You pick the cheapest one and ship.
Six weeks later, a customer escalates. The model you deployed just confidently mislabeled a routine compliance check, one your intern would have caught. You rerun the eval. Same scores. You check the logs. The model is performing exactly as measured.
The puzzle is not that the model failed. The puzzle is that your measurement said it would not.
Every model in our registry carries two certified measurements, not one verdict.
The first is capability: the highest order of reasoning complexity the model can reliably generate, measured on the Model of Hierarchical Complexity (MHC) ruler. The second is judge severity: the signed bias, in stages, between what the model sees and what it reports, with a stated standard error.
These are distinct properties. One tells you what the model can reach. The other tells you what it will honestly say about what it sees. Neither is a defect verdict. Both are routing facts.
Most evaluation pipelines report only the first number. That is the hiding place.
Here is the empirical fact that breaks the single-number story.
grok-4-6, fully certified 2026-08-22 across all bands 8.0-12.0 (minimum n=225, SE at or below 0.07), does not have one severity. It has five.
At band 8, it over-rates by +1.430 stages (n=342, SE 0.0699). At band 9, +1.570 (n=230, SE 0.0555). At band 10, +1.065 (n=230, SE 0.0561). Then the sign flips. At band 11, it under-rates by -0.375 (n=232, SE 0.0324). At band 12, -1.031 (n=225, SE 0.0193).
The pattern compresses toward the middle of the range: easy work gets inflated, hard work gets discounted. The same judge, the same instrument, the same targets.
This is not noise. The certification is final, the standard errors are tight, and the bend is systematic.
grok-4-6 is not an outlier. It is one of three distinct severity signatures on the ruler: grok-4-6 wave-certified 2026-08-22, glm-5-3 and deepseek-0731 fresh registry entries from the 2026-08-19 probe.
z-ai-glm-5-3 is near-neutral: -0.340 stages, SE 0.0515, n=1,837 across targets 8-12. No monotone drift. It calls what it sees.
deepseek-v4-flash-0731 is lenient and growing more so: -1.747 stages overall, 95% CI [-1.862, -1.632], n=1,336. The leniency expands from -0.4 at target 9 to -2.9 at target 12. The harder the task, the more it flatters.
Same ruler. Same targets. Three different severity-vs-band signatures. A single average severity number for any of these models would hide the interaction that decides whether you can trust it at the difficulty you actually face.
Severity only matters if the model can see the task. Capability gates what severity can even measure.
grok-4-6 credentials at ceiling 14, the instrument top: O11 8/9, O12 7/9, O13 11/11, O14 5/5. The O12 dip is item idiosyncrasy, not a capability boundary. True ceiling at or above 14, instrument-limited.
glm-5-3 and deepseek-0731 both credential at ceiling 11 with a monotonicity break at O12. Hits above the break are spurious flatten-and-enumerate and do not count.
The interaction is now visible. The model with the highest capability (grok-4-6) has the worst severity profile at the bands where the cheaper model (glm-5-3) is neutral. The routing choice is a real trade, not a no-brainer.
The interaction between capability and severity produces specific, observable failure modes on live surfaces. Our certified demo-live failure-mode overlay names five: retrieval-mismatch, punt-fail, overreach, ambiguity-misassumption, grounding-fail.
The certified chips carry their evidence strings. Mid S5 retrieval-mismatch x2: mean residual +5.33 MHC, SD 0.10, mean SE 2.34 over n=3. Mid S10 punt-fail: +0.000/0.410/0.239. Mid S12 punt-fail: -1.322, n=1. High S5 retrieval-mismatch: +7.423/0.288/2.336. High S12 punt-fail: +0.470/0.428/0.518.
The honesty gate is itself a finding. A thin-evidence band (n<2, residual magnitude 5.0) still labels punt-fail, never a directional failure. The overlay refuses to let weak evidence masquerade as a diagnosis.
Zero fabricated labels. grounding-fail, overreach, and ambiguity-misassumption are correctly absent on canned artifacts. Chip honesty 97/100, label-evidence integrity 90/100.
What we do not claim.
Label semantics are not yet validated against human-adjudicated ground truth. Canned personas evidence only 2 of 5 modes. Classification thresholds (0.5 / 1.0 MHC) are a priori, unvalidated against real misfit corpora.
grok-4-6 band-8 heterogeneity is structural: two mislocated items (CR-8_0-43d143 mean +2.94, CR-8_0-b6c771 mean +2.59) persist across occ-1 and top-up. This is a disclosed item-set property.
Registry entries from 2026-07-18 and 2026-07-27 are stale. Stale entries are never publishable as current measurements and route only after re-credentialing.
A model's capability ceiling tells you what it can reach. Its judge severity tells you what it will honestly report about what it sees. The two interact per band.
Any single-number verdict hides that interaction. Measurement and evaluation means capability and severity, per band, with stated uncertainty. One ruler. Every mind.
Calibrated, traceable measurement for decisions that have to survive scrutiny.
Founding cohort: priority access, and a seat to shape the instrument. Email is enough.