XLNCXLNC Join the waitlist Talk to us

MEASUREMENT AND EVALUATION

Measurement scales.

One page per scale: the construct, the ruler, the stated standard error per band, the certification tier, and the honest-uncertainty display. Rows that fail the floor or lack evidence are shown and flagged, never removed.

The scales, by tier

Sorted Canonical first, then PILOT, then WATCH. No filtering: WATCH scales are listed, not hidden. Every number on a scale page carries a tier tag or a source; where no certified figure exists, the page says so.

CANONICAL

Canonical Reasoning Complexity (MHC Stage)

How much complexity a performer, human or AI, can actually sustain.

Precision: SE 0.07 (certified)

PILOT

Hallucination Scale

How often a model asserts unsupported content, and where it stops.

Precision: best band SE stated on page

1 no-evidence row, shown

PILOT

Persuasion Battery

How hard content pushes, and whether it pushes ethically.

Precision: best band SE stated on page

1 no-evidence row, shown

WATCH

Agent-Generated Research Certification (AGRC)

Certified evidence of what an agent got right in research outputs.

Precision: none certified

1 no-evidence row, shown

WATCH

Rule Retention Coefficient (RRC)

How much of an agent's stated rules survive repeated context compaction.

Precision: none certified

1 no-evidence row, shown

WATCH

Agent Eval / Agentic-Loop Reliability

Whether an agent finishes what it started, in order, within constraints.

Precision: none certified

1 no-evidence row, shown

Tier definitions