XLNC Join the waitlist

Blog / 2026-09-02

Benchmaxxing is a measurement failure

The same model, the same benchmark, a different score every run. That instability is not noise in your results. It is your instrument telling you it was never measuring.

Run your favorite model against your favorite benchmark twice. Same model, same items, same harness. Two scores. Run it ten times and you will get a band, not a point. Every practitioner has seen this. Almost nobody asks the question it begs: if the number moves when nothing moved, what exactly was it measuring?

Here is the uncomfortable answer. Nothing. Not capability. Not performance. A benchmark score is a count of correct answers, and a count is not a measurement until every item contributes equally to the construct. They never do. A benchmark is an unmarked stick: it can show you that one thing is longer than another, sometimes, and it cannot tell you how much longer, and it quietly changes length when you look away.

This is the behavior we call benchmaxxing: optimizing the score instead of the construct. Not because engineers are careless, but because the score is the only instrument they were handed, and the score rewards its own optimization. Prompt tuning, item selection, decoding tricks, a favorable retry policy, and the number climbs while the capability underneath stays exactly where it was. The moment a number is treated as a goal, it stops being an observation. This is not a new insight about incentives. It is one of the oldest findings in the measurement literature, rediscovered every year in a new domain. The AI field is simply its latest renter.

Three failures follow, and each one alone should be disqualifying for a decision of consequence.

First, instability. A score that moves across runs with no change in the performer is an instrument with no calibration and no stated error. If a thermometer read 21, then 24, then 19 in a stable room, you would not argue about the room. You would replace the thermometer.

Second, incomparability. Raw scores have no common scale across benchmarks, versions, or systems. A 0.85 on one suite and a 0.72 on another are not comparable quantities, and neither is an amount of anything. They are rankings with decimal points attached, and rankings of different item mixes do not subtract, average, or trend.

Third, silence about error. A benchmark score arrives with no standard error, because the model that produces it does not generate one. The number looks precise to two decimal places and is silent about its own uncertainty, which is the exact opposite of what a decision needs.

NIST reached the same conclusion in its report on expanding the AI evaluation toolbox with statistical models (NIST AI 800-3, February 2026): "the intervals between LLM capabilities on the latent scale have meaning independent of the difficulties of the specific tasks selected into the benchmark. Raw benchmark scores do not have this property." Read that sentence twice. It is the national measurement institute of the United States telling the field that the difference between two raw scores is not a quantity. Not "is noisy." Not "should be interpreted with caution." Is not a quantity, because the intervals depend on which items happened to be selected.

The way out is old, published, and sitting in plain sight. Model item difficulty and performer capability on a single latent scale, so that ordinal right-and-wrong patterns convert into interval-scale estimates with a stated standard error at every value. This is Rasch conjoint measurement, and it is the reason a calibrated ruler behaves differently from an unmarked stick: its intervals mean the same thing everywhere on the scale, regardless of which items you happen to use, and every measurement carries its own uncertainty instead of hiding it.

That is the standard we hold our own work to, and it is the standard we think this field should demand of anyone selling it a number. On the AIM ruler, our current standard error is 0.07 per measurement. The method is published, not proprietary folklore: Chapter 18 of the Cambridge Handbook of Assessment Center Methods documents the unobtrusive measurement approach, and our calibration-validity program was registered on the Open Science Framework before data collection.

We hold that the AI evaluation community can do better than it currently does, and we intend to argue for it here, post by post. Not because benchmark teams lack talent, but because they were handed a broken instrument and asked to make decisions with it. The honest response to an unmarked stick is not to read it more carefully. It is to get a ruler.

This blog exists for that argument. Future posts will cover why we publish the error bar on every value, what it means to measure humans, AI agents, and robots on one calibrated scale, and what evaluation looks like when an agent acts continuously with no human watching it.

If a number is going to carry a decision, it should be a measurement. We think you agree, or you would not have read this far.

Join the waitlist if you want calibrated, traceable measurement for decisions that have to survive scrutiny.

Join the waitlist