A change arrives
A new task, a new model, a cost spike, or an incident. AIM does not invent the trigger. It is the ruler you pick up when the question is whether capability still holds after the change.
MEASUREMENT AND EVALUATION
Benchmarks score. AIM measures. One calibrated ruler, with the standard error stated at every value, so a moved number means a moved model and not a moved ruler.
Free to measure. Pay to certify. Every published value carries its stated standard error.
MEASURED OPINION, FROM PEOPLE WHO STAKE THEIR NAMES ON RIGOR
Dr. Robert Cialdini, NY Times Best Selling Author, Founder Cialdini Institute, Regents Professor Emeritus ASU
"I've always stressed the importance of ethics in persuasion, and Dr. Matt Barney's AI assessment tool brings unprecedented scientific rigor to this domain. I am optimistic that his method holds immense promise in proactively preventing the misuse of persuasion techniques, both by people and emerging technologies, and augmenting their long-term use correctly"
"Dr. Matt Barney is a blue chip of the consulting world. As an experienced business leader he has the instinct to know what works in practice and what doesn't; yet, he also understands the limits of personal knowledge and draws on science-based practices when needed. Those who work with Matt can be sure of one thing. His advice will not only be relevant but it will also be rigorous."
Professor John Antonakis, University of Lausanne; Editor-in-Chief, Leadership Quarterly
THE INDUSTRY DEFAULT
3 judges. 1 rating each. Averaged.
Uncalibrated. No stated uncertainty. No traceability.
AIM
Kaleidoscope: many viewpoints, not one prompt, so the bank measures the skill, not one model's way of asking
Hydra calibrated judges, inverted Computer-Adaptive Testing: items stop when the uncertainty target is met
Hexagon traceability: every measurement traces back to certified anchors, end to end
A measurement and evaluation instrument, not an average of opinions.
The left column is a vote. The right column is a ruler.
Benchmarks score tasks. They do not measure performers.
OPTIMAL PROCESS
STEP 1
Canonical MHC locates the capability a task, process, or agent minimally requires. That location becomes the business standard: the requirement every candidate performer is compared against, whether the candidate is an LLM, a person, or a machine. Quality first: performers may match on quality and still differ on cost, throughput, or latency. But if quality is poor, the rest do not matter.
STEP 2
Each candidate is measured against that standard with domain-specific measurement batteries, ideally on multiple samples. The output is process capability against the business standard: who clears the requirement, with what uncertainty, at what cost and speed.
Ready now: Canonical MHC task standards (CERTIFIED). In refinement: Hallucination and Persuasion scales.
MODEL LEADERBOARD
Most AI benchmarks give you a single continuous score. The Model of Hierarchical Complexity (MHC) gives you something different: an ordinal ladder of reasoning complexity. Each stage is a qualitatively different way of organizing information, not a point on a line. MHC is domain-general: it answers what is the most complex structure this system can process. Domain-specific batteries answer how well it actually does the task in this context. You need both: MHC tells you if the glass is big enough; the domain scale tells you how good the water tastes.
Before you pay to certify performance, you need to know: does the model clear the complexity floor your task requires? A massive engine does not win a race by itself, but you cannot win a complex track without a sufficient engine. Clear the stage your task requires, then move to certification. This is the domain-general MHC leaderboard; domain-specific leaderboards follow as those scales are certified, starting with Hallucination, then Persuasion. MHC qualifies; the domain scale certifies.
Domain-general, not domain-specific. This table qualifies; it does not rank, and no model is ranked against another here.
| Model | Complexity ceiling (MHC stage) | Cost | Latency (typical) | Reliability |
|---|---|---|---|---|
| z-ai-glm-5-3 | ▶ Stage 11 (certified with a disclosed condition: a consistency break above Stage 11, so higher hits do not count) | Low tier | 7.89 s | Pending |
| deepseek-v4-flash-0731 | ▶ Stage 11 (certified with a disclosed condition: a consistency break above Stage 11, so higher hits do not count) | About $0.0004 per rating | 9.86 s | Pending |
| grok-4-6 | ▶ Stage 14 (the top of the current instrument: the true ceiling is at least 14, limited by the instrument, not the model; a judging-side condition is disclosed) | High tier (input $2.27, output $6.80 per million tokens) | Pending | Pending |
| Model (family) | Complexity ceiling | Cost | Latency | Reliability | Status |
|---|---|---|---|---|---|
| Mistral family | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Alibaba family (235b) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Google family (flash) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Alibaba family (8-max understudy) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| xAI family (legacy grok) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| OpenAI open-weight family | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Meta family (70b) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Google family (pro) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| DeepSeek family (pro) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| DeepSeek family (flash, predecessor) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Moonshot family (k2) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| OpenAI family (gpt-5.x) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Moonshot family (k3) | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
| Fable family | ▶ Pending | Pending | Pending | Pending | Credentialing in progress |
▶ Certified on disk. ▶ Credentialing in progress, shown honestly as pending: no number is fabricated or implied. Reliability is pending for all three qualified models because none were in the most recent reliability panel; the gap is shown, not hidden.
Standard error note: the standard error for the stage a model can handle generically is stated in the registry of record for each certified model. It is carried there as a number, not drawn here as a confidence whisker.
Values verbatim from the XLNC registry of record, measured August 2026. Measure the glass before you judge the drink.
Uncertainty is certified to SE 0.10 per band on the newer scales, band 1 at SE 0.15. SE 0.07 is the in-validation target, labeled PILOT. Separately, the canonical MHC ruler holds a realized cert at SE 0.07 or better across bands 8.0 to 12.0 (Deming FINAL CERTIFY, 2026-08-22).
Method published: Chapter 3, Models, Measurement, and Metrology Extending the SI (Barney & Barney 2024, De Gruyter).
Traceability follows the Mari and Wilson metrology-psychometric hexagon framework: every construct specified from measurand to public value. Aligned with NIST AI 800-3 (February 2026), published across a 16-year research program, with a co-authored paper in preparation on MASEMS (Metrological Architecture for the Social, Educational, and Medical Sciences; Thornton, Barney and Fisher).
No AIM score is a black-box number; each one resolves to a specific item in a frozen bank.
the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.
NIST AI 800-3, February 2026. Quoted verbatim. The same claim runs once on this page: see the full citations or read the NIST Alignment.
AIM measures tighter than the industry's best standard and tighter than every familiar test. It is built like a precision instrument: the uncertainty is stated on every number, never buried.
The full 17-row comparison, with every source, lives on the ruler page.
An evals engineer does not ship a one-shot score. The work is a loop: trigger, design, build, measure, decide, monitor, iterate. AIM is the MEASURE layer of that loop: stated-SE capability measurement on an anchored ruler, plus a drift watch on the instrument. The rest of the loop stays yours.
A new task, a new model, a cost spike, or an incident. AIM does not invent the trigger. It is the ruler you pick up when the question is whether capability still holds after the change.
Write the task in the buyer's vocabulary and place the goal band on the AIM ruler. The standard is a location on an interval scale, with a stated standard error, not a pass-rate that moves when the item set moves.
You still build the eval set and the rater panel. AIM consumes evidence as performer, item, rater, and raw score. The judge is a measured facet of the instrument, not an invisible grader whose severity is averaged away.
This is AIM today. The engine returns a calibrated location plus a stated standard error on one anchored MHC ruler. Humans, AI agents, and robots sit on the same interval scale. No published value ships without its uncertainty. The measurement method is CERTIFIED; production control charts are not.
Compare the measure to the goal band with combined uncertainty: clearly above, indeterminate, or clearly below. Route, scaffold, gate, or redesign from that read. A bare scalar is not a decision.
Calibration vintage is recorded on every measure. Watch judge severity and performer location over time so a silent shift in the instrument is not mistaken for a shift in the work. Production control-chart operations are on the roadmap; the vintage stamp is already part of the measure.
Add failing cases, fix, and measure again on a stable instrument. The delta is then a real change in the performer or the work, not a moving benchmark. Repeatable harness, not a one-shot score.
Is this model capable enough for this task against a stated standard? AIM supplies the MEASURE layer: stated-SE capability on the anchored ruler, then a three-zone read against the goal band. Go, no-go, or scaffold with the uncertainty in view.
Which model clears the quality bar under real constraints? Measure quality on the AIM ruler with a stated SE. Keep cost, latency, and reliability as process metrics with their own spec bands. Gate on the ruler, then select. Do not treat a quality location on the ruler as a process-capability index.
Which parts of the work should humans keep, and which can an LLM hold, against a business standard? Measure both populations on the same ruler. Compare distributions, not one point estimate. Human performance is often heavy-tailed; a mean hides the tail you still need to staff.
Persistent agents and CI pipelines cannot open a dashboard at every gate. AIM is designed to be called as a headless API and MCP endpoint inside this loop: request a measurement, receive a calibrated result with a stated standard error and a calibration vintage, then decide. API-first, not API-only. The human UI remains the reporting and audit layer.
That surface is planned (F-API-MCP-1). It is not shipped. Do not write a tool call against it today. For early access, join below or write matt@xlnc.co.
Interval-scale measurement, construct maps, and judge severity diagnostics. The full method.
NIST AI 800-3 names the interval-scale property. See the verbatim quote and the scorecard.
Free to measure. Pay to certify. See tiers and delivery models.
One ruler. Every mind.
Decisions that have to survive scrutiny deserve measurement, not vibes.
Calibrated, traceable measurement for decisions that have to survive scrutiny. Measures, not vibes.
Founding cohort: priority access, and a seat to shape the instrument. Email is enough.
Not a product claim. We will write when there is something to measure. You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.
What happens next: a confirmation email (confirm to hold your place), a 20 minute call to scope your use case once confirmed, and an invitation as founding cohort slots open.
You are on the list. Check your inbox to confirm.