The Index / Data
All 150, with the text each score came from.
Every company named. Every verbatim message unit. Every score. Downloadable as CSV, free, no attribution required.
01Schema
Not yet published
Scoring runs October 2026. The data set goes live with the findings in November. The schema below is fixed now so the shape of it can't be adjusted after the numbers come in.
What each row will contain
| Field | Values |
|---|---|
| company | Named, in full |
| segment | One of six |
| revenue_usd_m | From the single named source |
| url | Full URL as captured |
| capture_date | YYYY-MM-DD |
| first_message_unit | Verbatim, including punctuation |
| position | negative_present / shortcut / positive_future / unclassifiable |
| level | informational / aspirational / transformational |
| band | dont_hurt_me / assertive / aggressive |
| upstream | 0–4 |
| swap_test | pass / fail |
| swap_competitor | The company substituted in |
| scorer | Initials |
| model_assisted | y / n |
| human_override | y / n |
| double_scored | y / n |
| adjudicated | y / n |
| notes | Free text |
02Alongside
Published alongside it
- The exclusion log: every company dropped from the frame, and why
- Reliability statistics for the double-scored subsample
- The model override rate
- The methodology, unchanged from its pre-registered version
03The reason
Why publish the whole thing
Because a finding you can't check is an opinion with a chart on it. If you disagree with a score, you will be able to see the exact sentence it was derived from and the rule that was applied to it — and if the rule was applied badly, you'll be able to prove that.
That is a real risk to take. It is also the only reason to believe any of this.