XLNCXLNC Join the waitlist Talk to us

MEASUREMENT SCALE 1

Canonical Reasoning Complexity (MHC Stage)

CANONICAL

How much complexity a performer, human or AI, can actually sustain.

All measurement scales · One calibrated ruler · The method

S1. Construct definition

CANONICAL

When you are about to trust a model with work that has real structure, you need to know how much complexity it can actually sustain, not how it did on a quiz. This scale measures the developmental order of hierarchical complexity a performer, human or AI, can sustain: the flagship ruler every other AIM scale re-anchors onto. It does not measure knowledge, style, or speed; a fluent wrong answer at a low stage is still a low-stage answer. A benchmark can tell you two models tied on a test form. This ruler tells you where each one sits on the complexity order, with the standard error stated at every band. Frozen bank, audit-ready.

Boundary. It does not measure knowledge, style, or speed; a fluent wrong answer at a low stage is still a low-stage answer.

Decision context. Feeds capability sufficiency, model selection, and work redesign decisions (the three-decisions block on the Science page).

S2. The Rasch/PCM ruler

Rasch family measurement model; calibrated item bank, pre-registered on OSF, anchored to human data.

Measurement frame. Units are logits on the MHC complexity order, with certified stage-boundary pins. A one-logit difference is a constant odds ratio everywhere on the ruler: intervals mean the same amount at every stage.

MHC reasoning complexity construct map card, Mari-Wilson hexagon, Logic-1 hierarchical
CERTIFIED MHC Reasoning Complexity - Logic-1 hierarchical. Canonical construct map with certified stage-boundary pins. The full comparison ruler with flagged rows lives at One Calibrated Ruler.

S3. Stated SE per band

Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.

BandBand labelStandard errorCertification implication
All published bandsCertified measurement-method standardSE 0.07CERTIFIED at SE 0.07 for the canonical MHC ruler (science.html, SE 0.07 standard card).
--BELOW FLOOR Per-band judge-evidence readout Model rows whose band severity SE exceeds the frozen quality floor render on the Science page select-ruler, flagged, never presented as a pick.--Shown and flagged, never removed.
--NO EVIDENCE Per-band no-evidence models A model with no per-band judge evidence at a band is listed under no evidence, never silently dropped. Both row classes render on the Science page select-ruler readout.--Shown and flagged, never removed.

S4. Certification tier

CANONICAL

Canonical: frozen bank, audit-ready.

Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.

S5. Honest-uncertainty display

Rows that fail the floor or lack evidence are shown and flagged, never removed.

The flagged rows below are fixtures standing in for the live select-ruler readout; the rendered rows with real model names are on the Science page.

The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.

S6. API and usage detail

Status. Live in the demo.

POST a response set against the MHC bank; receive a band location on the complexity order with the standard error stated at the value.

POST /score  {"scale": "mhc", "responses": [...]}
{"band": "S11", "location_logit": "<stated>", "se": "<stated at the value>"}

Rate and pricing tier per the pricing page; certification tier carries SE 0.10.

The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.

S7. Scientific references

  1. 2026 Barney, M., Wind, S., & Krishna, V. (2026). Using large language models to evaluate ethical persuasion text: A measurement modeling approach. IJATE, 13(1), 224-247. https://doi.org/10.21449/ijate.1788563
  2. 2016 Barney, M.F. & Fisher, W. (2016). Adaptive Measurement and Assessment. Annual Review of Organizational Psychology and Organizational Behavior, 3, 469-490. https://doi.org/10.1146/annurev-orgpsych-041015-062329
  3. OSF Calibration study pre-registration and audit artifacts (Open Science Framework), per the Science page research program.