MEASUREMENT SCALE 3
How hard content pushes, and whether it pushes ethically.
When your legal, sales, leadership, or medical content is drafted or delivered by a model, you need to know how hard it pushes and whether it pushes ethically, not whether one reviewer liked the tone. This battery measures the severity and effectiveness of influence attempts on calibrated scales, built for ethical-influence quality control. It does not measure whether content is truthful or compliant with a specific statute; it measures the influence mechanics themselves. A vibe check varies with the reviewer. This ruler does not: the same attempt lands in the same band, with the standard error stated. The 730-item battery is certified; the instrument is calibrated, evidence accumulating.
Boundary. It does not measure whether content is truthful or compliant with a specific statute; it measures the influence mechanics themselves.
Decision context. Feeds ethical-influence quality control for legal, sales, leadership, and medical content.
Partial Credit Model family; the 730-item battery is certified (canonical list, rank 3). Severity and effectiveness scales calibrated; evidence accumulating on deployment forms. PILOT.
Measurement frame. Units are logits on the influence-severity and influence-effectiveness orders; a one-logit difference is a constant odds ratio on the rated influence mechanics.
Every published measurement carries a stated standard error, stated per band, never as a single scale-wide average. Averages hide the floor.
| Band | Band label | Standard error | Certification implication |
|---|---|---|---|
| All bands | Per-band standard error | not yet published per band | PILOT: the battery is certified; per-band SE tables publish with the first calibration wave readout. No figure is asserted without a source. |
| -- | NO EVIDENCE Per-band SE table Not yet published for this battery; shown as absent rather than implied. The 730-item battery certification does not by itself publish per-band precision. | -- | Shown and flagged, never removed. |
PILOT: calibrated, evidence accumulating.
Tier as of vintage 2026-09-01. Tier changes are events; the promotion rules for WATCH scales are stated in the canonical scale-priority record.
Rows that fail the floor or lack evidence are shown and flagged, never removed.
The flagged row below is the honest-uncertainty state of this scale: certified bank, per-band precision readout pending.
The flagged rows render in the S3 table above with the shared .below-floor and .no-evidence classes, the same flagged treatment as the model-ratings readouts. Dropping a failing row is a defect class, not a style choice.
Status. Demo program committed for Sept 22 (canonical list, rank 3); external scoring follows the demo. PILOT.
Scoring signature will mirror the MHC endpoint: submit rated influence-attempt responses, receive scale locations with stated standard errors.
POST /score {"scale": "persuasion", "responses": [...]} (PILOT; not yet open)
{"band": "<band>", "location_logit": "<stated>", "se": "<stated per band>"} (shape only; endpoint not yet live)
The 730-item battery is the certified artifact; the scoring endpoint opens with the demo program.
The full API reference lives on the docs surface when it is funded (IA spec section 4); until then this block is the usage detail of record.