XLNCXLNC Join the waitlist
PROPOSED

Which models can you trust as instruments?

A continuously updated, independently produced reference dataset rating LLMs as measurement instruments and production components: capability by domain with stated standard error, judge severity, reliability, availability, cost, latency, and explicit fitness-for-use-as-an-instrument verdicts.

No ratings are published yet. The first dataset ships only after the measurement system behind it is certified.

What it is

A continuously updated, independently produced reference dataset that rates large language models as measurement instruments and production components. Not a leaderboard of benchmark scores. A measurement and evaluation product that treats every model the way a metrology lab treats a gauge: something to be calibrated, monitored, and given a stated uncertainty before you trust it in a decision path.

Each entry in the dataset describes one model, at one pinned version, across stated domains. Model providers change behavior silently; a rating pinned to a version and a date is the only kind that stays true.

What the dataset measures

Every numeric field carries a stated standard error or a stated observation window. No number without its uncertainty.

The shape of the free tier (illustrative format, not data)
Model (illustrative)Capability (MHC stage, SE)ReliabilityAvailabilityCostLatency
Frontier model A (illustrative)Band 12, SE 0.14High, 30-day window99.9%, 30-day windowLowFast
Frontier model B (illustrative)Band 11, SE 0.20Medium, 30-day window99.5%, 30-day windowMediumMedium
Efficient model C (illustrative)Band 10, SE 0.30Medium, 30-day window99.0%, 30-day windowLowFast

Illustrative format, not data. Instrument-fitness verdicts and per-domain severity stay behind the paid tiers, so the public table never carries the answer key.

The independence firewall, stated plainly

This product earns nothing from routing, hosting, or selling inference. No provider pays to be rated, pays to be removed, or pays for a better number. The dataset is funded by its subscribers, on the subscriber-pays model of the ratings agencies, and it is structurally separated from the measurement business it grew out of. We publish measurements about models, never verdicts about companies, and every fitness statement is pinned to a model version and a date. If a provider changes the model tomorrow, the dataset records the change; it does not generalize a verdict to a brand.

That separation is not a disclaimer. It is the feature that makes the numbers usable.

What you get

All packaging below is PROPOSED. Prices are PROPOSED and untested; willingness to pay has not been measured.

Honest cadence

Certified ratings refresh quarterly. Telemetry updates monthly. Models churn faster than quarterly, and we state that limitation in the methodology rather than hide it. The trade is freshness against precision, and this product sells precision: a slightly older number with a stated standard error beats a fresh number with no uncertainty attached.

Why nothing is published yet

The first dataset is in production, not in public. Nothing ships until our measurement-system analysis clears: gauge repeatability and reproducibility on the scoring instrument first, every capability, severity, and fitness number certified before it is named. That claim hold is the same discipline we apply to every score we publish about people. It applies to models too.

Join the waitlist and you will hear the moment the first certified vintage is dated.

Get the first certified vintage

Join the waitlist and we will tell you when the first certified dataset is dated. Free summary first; subscriptions and certified reports follow the claim-hold gate.