XLNCXLNC Join the waitlist Talk to us

Blog

Ground truth is a faith, not a measurement

When you want to know if an AI system is right, you ask a human, and the human's rating is called ground truth. But the human reference itself carries systematic, measurable bias that changes with the capability of the thing being measured, and almost nobody has ever checked.

2026-09-08 · measurement

Here is a belief so common it goes unquestioned: when you want to know if an AI system is right, you ask a human. The human's rating is the "ground truth." The machine is judged against it. Case closed.

But what if the ground truth moves?

Not philosophically. Measurably. What if the human reference itself carries systematic bias that changes with the capability of the thing being measured, and what if nobody has ever checked?

The weakness of human ground truth

Psychometrics did not invent the many-facet Rasch model because human ratings were reliable. It invented the model because they were not.

Myford and Wolfe (2003, 2004) catalogued the taxonomy: severity and leniency, halo, central tendency, restriction of range. These are not noise. They are "rater effects," systematic and estimable, which is why Linacre (1989) built the MFRM to parameterize them. Engelhard and Wind (2018) devote an entire monograph to "rater-mediated assessment," a subfield that exists precisely because raters are fallible instruments requiring calibration. Eckes (2015) teaches the same taxonomy as standard curriculum.

Human raters drift. They vary with expertise. They need training, monitoring, and recalibration (Linacre et al., 1994; Myford & Wolfe, 2009; Noh & Matore, 2020; Sampson et al., 2024; Yan & Chuang, 2023, as cited in Thornton, Barney, & Fisher, in preparation). And humans cannot be re-run deterministically. The same rater, the same response, a different day, a different mood, a different lunch: the observation changes.

This is not a scandal. It is a known property of human judgment, and psychometrics built machinery to handle it. The scandal is that the AI evaluation field largely ignores the machinery while still treating the human as the fixed point.

The estimable machine

An LLM judge is different in kind, not just in degree.

You can fix the prompt. You can fix the seed. You can re-run the same response through the same judge a thousand times and get the same output, or a distribution you can characterize. That means bias is estimable. Once estimated, it is correctable. The correction has a known ceiling: you cannot correct what you cannot measure, but you can measure an LLM's bias with arbitrary precision given enough replications.

The economics are stark. In our pilot-stage cost ledger, certifying a judge costs $0.10. Skipping certification risks a $7.74 to $15 wave re-run plus bank contamination. The Phase 1 probe ran 480 cells for $0.1224 total. Deterministic re-runs at negligible marginal cost are not a hope. They are telemetry.

The sign flip nobody was looking for

Here is what happens when you actually measure an LLM judge across capability levels.

In the XLNC Phase 1 pilot probe, Nemotron Nano 30B, a frozen W2R2 bank, 480 cells: on low-capability weak responses, the judge is lenient by +2.29 raw points (n=12). On high-capability weak responses, lenient by +2.60 (n=15). On high-capability strong responses, severe by -0.46 (n=24).

The verbatim finding: "Nano is systematically LENIENT on weak responses (+2.3 to +2.6 raw points) and neutral-to-severe on strong (-0.46 at high band). An omnibus correction would leave a large residual exactly on the weak tail that decides band assignment."

The bias changes sign. Lenient on the weak, severe on the strong. A single-number correction, the kind you would apply if you treated the judge as a fixed reference, would average the +2.6 and the -0.46 into a small residual and call it calibrated. The error would remain exactly where it matters most: on the boundary between weak and strong, where band assignment is decided.

This is not a theoretical risk. A prior published XLNC finding showed the same pattern in a different model: over-rating at bands 8 through 10 (+1.430, +1.570, +1.065), flipping to under-rating at bands 11 and 12 (-0.375, -1.031), with per-band sample sizes and standard errors.

The honest asymmetry

Here is what we have not done. We have not measured capability-level-varying bias in human raters under our design.

Our Phase 1 ledger measures LLM judges against incumbent-LLM consensus. There is no human rater arm. The WE2 manuscript describes a planned comparison, "a subset of approximately 5,000 responses that receive both human and LLM scores," but states these are "predicted outcomes ... not completed empirical findings."

This asymmetry is itself a finding. The field assumes humans are the reference. We have shown the reference can be measured, and that when you measure it, its bias changes sign across the very capability levels you care about. The assumption that humans are stable is untested in our design, and largely untested anywhere.

That is a limitation of our work. It is also the point.

Falsifiability beats faith

Rasch measurement with specific objectivity gives you a falsifiable claim: the same response, presented to different calibrated judges, yields the same measure within a stated standard error. If it does not, the model is wrong, and you can see where. The claim can fail. That is what makes it measurement.

"Ground truth" cannot fail. It is defined as true. You cannot recalibrate it because there is no calibration target above it. You cannot estimate its bias because it is the baseline against which bias is defined. It is not a measurement. It is a faith.

The MFRM canon (Myford & Wolfe, 2003, 2004; Linacre, 1989; Engelhard & Wind, 2018; Eckes, 2015) exists because human raters are measurable instruments with estimable biases. The metrological canon (Pendrill, 2019; Pendrill & Fisher, 2015; Mari & Wilson, 2014) exists because measurement requires traceability and uncertainty budgets. NIST AI 800-3 stays statistical and never crosses into metrology.

XLNC/AIM stands on both canons. The MASEMS architecture (Thornton, Barney, & Fisher, in preparation) extends metrological traceability to social and educational measurement. Our prior work has quantified LLM-as-judge rater bias at omnibus level (Barney, Wind, & Krishna, 2026; Barney & Barney, 2024).

What to do instead

Stop asking "is it aligned with ground truth?" Start asking three questions.

What is the judge's bias at each capability level, with what standard error? Is the bias stable across replications, or does it drift? And what is the uncertainty budget from raw response to reported score?

A human rater cannot answer these questions deterministically. An LLM judge can. That is not a reason to trust the machine blindly. It is a reason to measure the machine, correct what you find, and stop pretending the human alternative was ever measured at all.

References

Barney, M., & Barney, F. (2024). Transdisciplinary measurement through AI: hybrid metrology and psychometrics powered by large language models. De Gruyter. https://doi.org/10.1515/9783111036496-003

Barney, M. F., Wind, S. A., & Krishna, A. (2026). Quantifying LLM-as-judge rater bias with the many-facet Rasch model. International Journal of Assessment Tools in Education. https://doi.org/10.21449/ijate.1788563

Eckes, T. (2015). Introduction to many-facet Rasch measurement: Analyzing and evaluating rater-mediated assessments (2nd revised and updated ed.). Peter Lang.

Engelhard, G., Jr., & Wind, S. A. (2018). Invariant measurement with raters and rating scales: Rasch models for rater-mediated assessments. Routledge.

Keller, A. J., Kwegyir-Aggrey, K., Steed, R., Rao, A. K., Sharp, J. L., & Bergman, A. S. (2026). Expanding the AI Evaluation Toolbox with Statistical Models (NIST AI 800-3). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.800-3

Linacre, J. M. (1989). Many-facet Rasch measurement. MESA Press.

Mari, L., & Wilson, M. (2014). Measurement across the sciences: Rasch models for measurement. Measurement, 51, 315-327.

Myford, C. M., & Wolfe, E. W. (2003). Detecting and measuring rater effects using many-facet Rasch measurement: Part I. Journal of Applied Measurement, 4(4), 386-422.

Myford, C. M., & Wolfe, E. W. (2004). Detecting and measuring rater effects using many-facet Rasch measurement: Part II. Journal of Applied Measurement, 5(2), 189-227.

Pendrill, L. R. (2019). Quality Assured Measurement: Unification across Social and Physical Sciences. Springer. https://doi.org/10.1007/978-3-030-28695-8

Pendrill, L., & Fisher, W. P., Jr. (2015). Counting and quantification: Comparing psychometric and metrological measurement. Measurement, 71, 46-55.

Rasch, G. (1960). Studies in mathematical psychology: I. Probabilistic models for some intelligence and attainment tests. Nielsen & Lydiche.

Thornton, A. M. A., Barney, M., & Fisher, W. P., Jr. (in preparation). Metrological Architecture for the Social, Educational, and Medical Sciences (MASEMS): From measurement chaos to traceability.

Calibrated, traceable measurement for decisions that have to survive scrutiny.

Founding cohort: priority access, and a seat to shape the instrument. Email is enough.

Join the waitlist