MEASUREMENT AND EVALUATION
One page per scale: the construct, the ruler, the stated standard error per band, the certification tier, and the honest-uncertainty display. Rows that fail the floor or lack evidence are shown and flagged, never removed.
Sorted Canonical first, then PILOT, then WATCH. No filtering: WATCH scales are listed, not hidden. Every number on a scale page carries a tier tag or a source; where no certified figure exists, the page says so.
How much complexity a performer, human or AI, can actually sustain.
Precision: SE 0.07 (certified)
How often a model asserts unsupported content, and where it stops.
Precision: best band SE stated on page
1 no-evidence row, shown
How hard content pushes, and whether it pushes ethically.
Precision: best band SE stated on page
1 no-evidence row, shown
Certified evidence of what an agent got right in research outputs.
Precision: none certified
1 no-evidence row, shown
How much of an agent's stated rules survive repeated context compaction.
Precision: none certified
1 no-evidence row, shown
Whether an agent finishes what it started, in order, within constraints.
Precision: none certified
1 no-evidence row, shown