XLNCXLNC Join the waitlist Talk to us
CERTIFIED

MEASUREMENT AND EVALUATION

Your benchmark score moved. Did the model?

Benchmarks score. AIM measures. One calibrated ruler, with the standard error stated at every value, so a moved number means a moved model and not a moved ruler.

Free to measure. Pay to certify. Every published value carries its stated standard error.

A measured result with its uncertainty intervals meets the required minimum on the AIM ruler A horizontal logit ruler for Reasoning Complexity (MHC Stage) from 0.0 to 8.0. A point estimate at 4.0 carries four nested 95 percent confidence bands at standard errors 0.30, 0.20, 0.14, and 0.10. An ink tick at the required minimum business standard falls inside the widest bands, and a status tag reads Requirement met. Reasoning Complexity (MHC Stage), logit scale. Bands: 95 percent intervals at SE 0.30, 0.20, 0.14, 0.10 (widest to tightest). Requirement met 0.0 4.0 6.0 7.0
A measured result with its stated uncertainty, set against the required minimum. The 95 percent interval contains the standard, so the requirement is met at that confidence. Tighter standard errors narrow the interval: measurement and evaluation sharpen the decision.

MEASURED OPINION, FROM PEOPLE WHO STAKE THEIR NAMES ON RIGOR

"Dr. Barney's multi-disciplinary approach draws from different business disciplines to develop an integrated model for value creation. It will serve as a guide to leaders in creating value on a sustainable basis."

N. R. Narayana Murthy, Founder and Chairman Emeritus, Infosys

"Dr. Matt Barney is a blue chip of the consulting world. As an experienced business leader he has the instinct to know what works in practice and what doesn't; yet, he also understands the limits of personal knowledge and draws on science-based practices when needed. Those who work with Matt can be sure of one thing. His advice will not only be relevant but it will also be rigorous."

Professor John Antonakis, University of Lausanne; Editor-in-Chief, Leadership Quarterly

"I can't say enough good things about Matt..."

Nathalie Salles-Olivier, Senior Director, Splunk; former Global Head of Live Learning, Facebook/Meta

Joel DiGirolamo, VP of Coaching Science, International Coaching Federation

Benchmarks score. AIM measures.

THE INDUSTRY DEFAULT

3 judges. 1 rating each. Averaged.

Uncalibrated. No stated uncertainty. No traceability.

AIM

Kaleidoscope relevance-gated diversity

Hydra calibrated judges, inverted CAT: items stop when the SE target is met

Hexagon traceability: six-sided provenance

A measurement and evaluation instrument, not an average of opinions.

The left column is a vote. The right column is a ruler.

Benchmarks score. AIM measures.

Benchmarks score tasks. They do not measure performers.

  • A 0.85 today is not a 0.85 tomorrow.
  • No two systems share a scale.

OPTIMAL PROCESS

What does this decision actually require?

STEP 1

Locate the task on the ruler

Canonical MHC locates the capability a task, process, or agent minimally requires. That location becomes the business standard: the requirement every candidate performer is compared against, whether the candidate is an LLM, a person, or a machine. Quality first: performers may match on quality and still differ on cost, throughput, or latency. But if quality is poor, the rest do not matter.

STEP 2

Measure the performers

Each candidate is measured against that standard with domain-specific measurement batteries, ideally on multiple samples. The output is process capability against the business standard: who clears the requirement, with what uncertainty, at what cost and speed.

Ready now: Canonical MHC task standards (CERTIFIED). In refinement: Hallucination and Persuasion scales.

Certified precision, published method.

CERTIFIED

SE 0.07 standard

The measurement method is certified at SE 0.07 for the canonical MHC ruler. The 0.07 target for newer scales is in validation. PILOT.

CERTIFIED

Rasch conjoint measurement

Method published: Chapter 18, Cambridge Handbook of Technology and Employee Behavior.

CERTIFIED

Traceability

Traceability follows the Mari and Wilson metrology-psychometric hexagon framework: every construct specified from measurand to public value. Aligned with NIST AI 800-3 (February 2026), published across a 16-year research program, with a co-authored paper in preparation on MASEMS (Metasystematic Assessment of Systematic Evaluations and Meta-Syntheses; Thornton, Barney and Fisher).

the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property.

NIST AI 800-3, February 2026. Quoted verbatim. The same claim runs once on this page: see the full citations or read the NIST Alignment.

One calibrated ruler

AIM sits tighter than the 10:1 gold decision floor, the 4:1 good floor, and every familiar test, on the same relative-precision frame as physical metrology instruments.

The full 17-row comparison, with every source, lives on the ruler page.

Open the one calibrated ruler

How AIM sits in the AI-evals loop

An evals engineer does not ship a one-shot score. The work is a loop: trigger, design, build, measure, decide, monitor, iterate. AIM is the MEASURE layer of that loop: stated-SE capability measurement on an anchored ruler, plus a drift watch on the instrument. The rest of the loop stays yours.

01 Trigger

A change arrives

A new task, a new model, a cost spike, or an incident. AIM does not invent the trigger. It is the ruler you pick up when the question is whether capability still holds after the change.

02 Design

Set the task and the standard

Write the task in the buyer's vocabulary and place the goal band on the AIM ruler. The standard is a location on an interval scale, with a stated standard error, not a pass-rate that moves when the item set moves.

03 Build / sample

Curate evidence and raters

You still build the eval set and the rater panel. AIM consumes evidence as performer, item, rater, and raw score. The judge is a measured facet of the instrument, not an invisible grader whose severity is averaged away.

04 Measure

The anchored ruler

This is AIM today. The engine returns a calibrated location plus a stated standard error on one anchored MHC ruler. Humans, AI agents, and robots sit on the same interval scale. No published value ships without its uncertainty. The measurement method is CERTIFIED; production control charts are not.

05 Decide

Read the band, then act

Compare the measure to the goal band with combined uncertainty: clearly above, indeterminate, or clearly below. Route, scaffold, gate, or redesign from that read. A bare scalar is not a decision.

06 Monitor

Drift watch

Calibration vintage is recorded on every measure. Watch judge severity and performer location over time so a silent shift in the instrument is not mistaken for a shift in the work. Production control-chart operations are on the roadmap; the vintage stamp is already part of the measure.

07 Iterate

Re-run the same ruler

Add failing cases, fix, and measure again on a stable instrument. The delta is then a real change in the performer or the work, not a moving benchmark. Repeatable harness, not a one-shot score.

Three decisions the loop is for

Use case 1

Capability sufficiency

Is this model capable enough for this task against a stated standard? AIM supplies the MEASURE layer: stated-SE capability on the anchored ruler, then a three-zone read against the goal band. Go, no-go, or scaffold with the uncertainty in view.

Use case 2

Model selection

Which model clears the quality bar under real constraints? Measure quality on the AIM ruler with a stated SE. Keep cost, latency, and reliability as process metrics with their own spec bands. Gate on the ruler, then select. Do not treat a quality location on the ruler as a process-capability index.

Use case 3

Work redesign

Which parts of the work should humans keep, and which can an LLM hold, against a business standard? Measure both populations on the same ruler. Compare distributions, not one point estimate. Human performance is often heavy-tailed; a mean hides the tail you still need to staff.

PLANNED

Headless API and MCP, not yet callable

Persistent agents and CI pipelines cannot open a dashboard at every gate. AIM is designed to be called as a headless API and MCP endpoint inside this loop: request a measurement, receive a calibrated result with a stated standard error and a calibration vintage, then decide. API-first, not API-only. The human UI remains the reporting and audit layer.

That surface is planned (F-API-MCP-1). It is not shipped. Do not write a tool call against it today. For early access, join below or write matt@xlnc.co.

Go deeper

The Science

Interval-scale measurement, construct maps, and judge severity diagnostics. The full method.

NIST Alignment

NIST AI 800-3 names the interval-scale property. See the verbatim quote and the scorecard.

Pricing

Free to measure. Pay to certify. See tiers and delivery models.

One ruler. Every mind.

Decisions that have to survive scrutiny deserve measurement, not vibes.

Join the waitlist

Calibrated, traceable measurement for decisions that have to survive scrutiny. Measures, not vibes.

Founding cohort: priority access, and a seat to shape the instrument. Email is enough.

Not a product claim. We will write when there is something to measure. You will get a short note from Matt when a measurement slot opens. Nothing else, no drip sequence.

What happens next: a confirmation within 1 business day, a 20 minute call to scope your use case, and an invitation as founding cohort slots open.