Every tier maps to a standard error target. Start free, then buy the precision your decision requires.
Free to measure. Pay to certify.
One calibrated ruler measures every performer: human, AI agent, robot. One interval-scale metric. Stated uncertainty.
Score anything for free. Every score comes with a standard error on an interval scale, the same property NIST AI 800-3 names directly: "the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property."
When you need tighter uncertainty, you move down the ladder. Each tier below is PROPOSED pricing under active test; it is not final and not yet contractual.
PROPOSED tiers:
The ceiling is SE 0.07, demonstrated on the AIM ruler and available in limited lighthouse slots. PROPOSED: allocation is capped while we validate throughput at that precision. If your decision requires publishable measurement-grade certainty, request a slot now or wait for capacity.
Every analysis runs on a unidimensional Rasch hallucination construct, so scores mean the same thing across performers and over time. Our Coaching Supervision stack is Rasch-analyzed end to end. Chapter 18 of the Cambridge Handbook covers our unobtrusive measurement approach: no prompts engineered to catch errors, no behavior changed by being watched.
Start with free scoring. See your number and its uncertainty before you spend anything. Move down the ladder only when the decision in front of you demands it.
Button: Start Free Scoring
Supporting line: No card required. Your first interval-scale score and its standard error are free. When you are ready to certify rather than screen, request a Defensible Tier quote.
| Tier | Standard Error | Use case | Price | Status |
|---|---|---|---|---|
| Screening / Coaching | SE 0.30 | Developmental feedback, coaching | $150 to $400 | PROPOSED |
| Default | SE 0.20 | Standard analysis | $1,500 per analysis | PROPOSED |
| Defensible | SE 0.14 | High-stakes, audit-ready | $3,000 to $5,000 | PROPOSED |
| Certification | SE 0.10 | Formal certification programs | $250K to $750K per year | PROPOSED |
| Ceiling (lighthouse slots, limited) | SE 0.07 | Measurement-grade certification | On request | PROPOSED |
One frame of reference. Conditional standard error at the decision point, in logits. Metrology gold, good, and Six Sigma sit on the same axis as the familiar tests.
XLNC / AIM 0.07 design target (PILOT) 10:1 gold 0.10 4:1 good 0.25 Six Sigma 0.33 (Westgard)
Conditional SE at the decision point (logits, k=1). Tight on the left, wide on the right.
What this means: every mark is the same quantity, conditional standard error at the decision point, so a lower bar is a tighter measurement. XLNC 0.07 sits tighter than even the metrology gold 10:1 reference.
GMAT is held: no like-for-like published precision figure exists, so it is not plotted.
Physical instruments (vernier 0.02 mm, micrometer 0.01 mm, gauge blocks, SI metre) define minimum precision as a stated floor in their native units. The metrology discipline, a stated decision-quality uncertainty floor, is what the gold, good, and Six Sigma anchors carry onto this ruler. AIM is not a NIST or ISO standard; the claim is that the discipline is the same. Test names are trademarks of their owners; XLNC is not affiliated with, endorsed by, or sponsored by any listed organization. Crosswalk: logit = 2 x (SE/SD); within-sample SD is about 1 MHC stage, about 2 logits (explicit modeling assumption).
Certified 07-23 figures only. Axis measurand is conditional standard error at the decision or cut point, in logits, k=1. Internal-consistency alpha, IRT marginal reliability, split-half, and composite reliability are not on this ruler. SAT and WAIS-IV use the certified 07-23 test-retest fallback. GMAT is held because no like-for-like published precision figure exists. This is not a claim that XLNC is better than any named test.
| Instrument | Conditional SE (logits) | What is plotted |
|---|---|---|
| XLNC / AIM | 0.07 | 0.07 is the XLNC / AIM design target in validation. PILOT. |
| ACT | 0.34 | conditional SE at the cut |
| WAIS-IV | 0.40 | test-retest |
| SAT | 0.53 | test-retest |
| GRE Quant at 150 | 0.55 | conditional SE at the cut |
| GRE Verbal at 150 | 0.64 | conditional SE at the cut |
| USMLE Step 3 | 0.67 | global SEM on a known score SD |
| USMLE Step 2 CK | 0.80 | global SEM on a known score SD |