A continuously updated, independently produced reference dataset rating LLMs as measurement instruments and production components: capability by domain with stated standard error, judge severity, reliability, availability, cost, latency, and explicit fitness-for-use-as-an-instrument verdicts.
No ratings are published yet. The first dataset ships only after the measurement system behind it is certified.
A continuously updated, independently produced reference dataset that rates large language models as measurement instruments and production components. Not a leaderboard of benchmark scores. A measurement and evaluation product that treats every model the way a metrology lab treats a gauge: something to be calibrated, monitored, and given a stated uncertainty before you trust it in a decision path.
Each entry in the dataset describes one model, at one pinned version, across stated domains. Model providers change behavior silently; a rating pinned to a version and a date is the only kind that stays true.
Every numeric field carries a stated standard error or a stated observation window. No number without its uncertainty.
| Model (illustrative) | Capability (MHC stage, SE) | Reliability | Availability | Cost | Latency |
|---|---|---|---|---|---|
| Frontier model A (illustrative) | Band 12, SE 0.14 | High, 30-day window | 99.9%, 30-day window | Low | Fast |
| Frontier model B (illustrative) | Band 11, SE 0.20 | Medium, 30-day window | 99.5%, 30-day window | Medium | Medium |
| Efficient model C (illustrative) | Band 10, SE 0.30 | Medium, 30-day window | 99.0%, 30-day window | Low | Fast |
Illustrative format, not data. Instrument-fitness verdicts and per-domain severity stay behind the paid tiers, so the public table never carries the answer key.
This product earns nothing from routing, hosting, or selling inference. No provider pays to be rated, pays to be removed, or pays for a better number. The dataset is funded by its subscribers, on the subscriber-pays model of the ratings agencies, and it is structurally separated from the measurement business it grew out of. We publish measurements about models, never verdicts about companies, and every fitness statement is pinned to a model version and a date. If a provider changes the model tomorrow, the dataset records the change; it does not generalize a verdict to a brand.
That separation is not a disclaimer. It is the feature that makes the numbers usable.
All packaging below is PROPOSED. Prices are PROPOSED and untested; willingness to pay has not been measured.
Certified ratings refresh quarterly. Telemetry updates monthly. Models churn faster than quarterly, and we state that limitation in the methodology rather than hide it. The trade is freshness against precision, and this product sells precision: a slightly older number with a stated standard error beats a fresh number with no uncertainty attached.
The first dataset is in production, not in public. Nothing ships until our measurement-system analysis clears: gauge repeatability and reproducibility on the scoring instrument first, every capability, severity, and fitness number certified before it is named. That claim hold is the same discipline we apply to every score we publish about people. It applies to models too.
Join the waitlist and you will hear the moment the first certified vintage is dated.
Join the waitlist and we will tell you when the first certified dataset is dated. Free summary first; subscriptions and certified reports follow the claim-hold gate.