XLNCXLNC

Reference definitions for the measurement model behind AIM. Each construct has a stable anchor; link to the anchor, not the page.

Rasch/PCM conjoint measurement#

Anchor: `#rasch-pcm-conjoint-measurement`

AIM scales are Rasch-family measurement models: the dichotomous Rasch model where items are scored right or wrong, and the Partial Credit Model (PCM) where items have ordered performance categories. Conjoint measurement means performer locations and item difficulties are estimated on one shared logit scale, so a performer's location means the same thing regardless of which subset of items they answered.

Properties that follow, and that benchmarks do not have:

  • Interval scale. A one-logit difference is a constant odds ratio everywhere on the ruler, so distances and averages are meaningful.
  • Specific objectivity. Comparisons between performers do not depend on which items were administered, and comparisons between items do not depend on which performers answered them.
  • Stated uncertainty. Every estimated location carries a standard error computed at that location, not a scale-wide average.

The method is published: Chapter 18, Cambridge Handbook of Technology and Employee Behavior (Barney 2019), and the 2016 Annual Review paper on adaptive measurement. See Scientific references.

Stated standard error per band#

Anchor: `#stated-standard-error-per-band`

Standard errors are reported per band, never as one scale-wide average. Averages hide the floor: an instrument can be tight in the middle of its range and loose at the decision point, and a single mean SE would report it as fine.

Two worked examples from the current instrument set:

  • Canonical MHC ruler: the measurement method is CERTIFIED at SE 0.07. The 0.07 target for newer scales is in validation, labeled PILOT wherever it appears.
  • Hallucination scale (two-layer cumulative construct, PILOT): interior bands carry a stated design precision of SE 0.10 per band; the first band carries SE 0.15. Both figures are design values, not yet certified, and are labeled PILOT everywhere they render.

Honest-uncertainty display rule, binding on every AIM surface: rows that fail the floor or lack evidence are shown and flagged, never removed.

Certification tiers#

Anchor: `#certification-tiers`

Every scale and every claim carries one of three tiers, rendered from a shared lexicon so the marketing site and this docs surface cannot drift:

  • Canonical: frozen bank, audit-ready.
  • PILOT: calibrated, evidence accumulating.
  • WATCH: provisional; decision-use prohibited for high-stakes calls.

A tier is stated as of a vintage date. Tier changes are events: promotion and demotion are recorded with the evidence that drove them.

The two-layer certification doctrine#

Anchor: `#two-layer-certification-doctrine`

Two-layer cumulative scales (the hallucination scale is the current instance) are certified in two layers, because the floor band and the interior bands carry different precision:

  • Interior bands (band 2 and above): stated design precision SE 0.10 per band.
  • Floor band (band 1): stated design precision SE 0.15.

Both figures are PILOT: stated design precision, not yet certified. The doctrine exists so that a looser floor band is declared up front rather than averaged away. Where these figures render on the site, they carry the PILOT pill and the words "not yet certified."

Renderer honesty#

Anchor: `#renderer-honesty`

Wherever a scale or model is listed, failing states render. Below-floor rows render with muted ink on a warm wash; no-evidence rows render flagged. Dropping a failing row is treated as a defect class, not a style choice. This rule is enforced in the site generator and checked on every rebuild.