Note
What makes a prediction testable
Four things have to be present. Most public predictions have two or three, which is why so few of them ever get marked right or wrong.
A prediction can be settled when four things are present. Most public predictions have two or three, which is why so few ever get marked right or wrong.
- A measurable outcome. Something with a value, or that definitely happens or doesn’t. “Bitcoin’s price” qualifies. “The vibe shifts” does not.
- A threshold. The number separating right from wrong. “Bitcoin goes up” has an outcome and no threshold, so it is unfalsifiable by a cent.
- A deadline. The date the question gets asked. Without one a claim is permanently pending and its author is never wrong, only early.
- A source that publishes the answer. The one everybody forgets. Somebody has to report the value, publicly, on a schedule, in a form both sides would have accepted before knowing the result.
The fourth is where most claims die
“We’ll triple our compute again next year” has an outcome, a threshold and a deadline. It fails anyway, because no published series reports OpenAI’s compute. Same with “we are going to again fail in 2026 to meet demand” — dated, specific, and nothing on earth measures met demand.
Against that, Nate Silver’s “Mike Johnson remains Speaker of the House through 11/3/2026.” Outcome: who holds the office. Threshold: it is him or it is not. Deadline: written into the sentence. Source: the House. Nothing left to argue about afterwards.
Two traps that quietly change the question
When. “Hits $200k by December” and “is above $200k in December” are different claims, and the gap is where most arguments end up. One is satisfied by a single tick. The other needs the price to be there on the day.
Which print. GDP, CPI and unemployment are revised after release. If a claim rests on one, the version that counts has to be fixed in advance — advance estimate, second, third — or the same prediction is right in October and wrong in December.
What to do with the rest
Fifty of the 101 records here fail this test and are kept anyway. A statement that cannot be scored is still worth having with its date and source attached: most of what anyone says in public is unscoreable, and an archive holding only the tidy half would misrepresent what people actually claim.
They are simply never given a verdict. The failure to avoid is not keeping them — it is inventing a threshold the speaker never gave so the archive can show a number. A precise-looking figure resting on a made-up criterion is worse than an honest blank.
The full rules, including how outcomes are decided and how to challenge one, are on the methodology page.
Every figure above is a count from the record on August 29, 2026, and the archive is exportable if you want to check the arithmetic yourself.
Also here: People who write predictions get scored. People who say them don't. · The goalposts problem · What a 96% accuracy score actually measures · How a track record lies without a single false statement · How to check whether someone actually predicted what they say they did · Sometimes the honest verdict is that nobody can tell · Everyone predicts in January and finds out in December