Note

The goalposts problem

Everyone knows to screenshot the claim. Almost nobody writes down what would settle it, which is the half that gets quietly rewritten.

Everyone knows to screenshot the claim. Almost nobody writes down what would settle it, and that is the half that gets rewritten.

The move is familiar. Someone predicts a thing. The thing roughly doesn’t happen. And the standard turns out to have been different all along — they meant the trend, not the number; they meant broadly, not by that date; the point was directional. None of it is necessarily dishonest. Memory is genuinely reconstructive, and people are sincerely confident about what they meant.

It works because the claim and the test for the claim get separated. The claim is public and dated. The test lives in someone’s head, where it is free to move.

What a screenshot proves, and what it can’t

A screenshot proves wording. That is not nothing — deleted posts are real, and so is quiet editing. But wording was rarely the disputed part. The dispute is almost always about what would have counted, and a screenshot is silent on that.

So the useful thing to record is not just the sentence. It is the sentence and the condition that decides it, fixed at the same moment, before anyone knows the answer.

On this record that condition is written down when the words support one, hashed together with the quote, and committed to a public blockchain in a batch. Not because a blockchain makes anything true — it doesn’t, and anyone claiming otherwise is selling something — but because it makes a later edit detectable by someone who dislikes us and has no special access. That is the only property worth having.

The uncomfortable half

Fixing the test in advance cuts both ways, and mostly it cuts against the person keeping score rather than the person being scored.

Locked tests resolve in ways that feel wrong. A forecast can be right for entirely lucky reasons and the test still says met. Someone can be substantially correct about the world and miss the threshold by a rounding error. The temptation afterwards is to explain why this particular result should be read with context — which is exactly the move the lock exists to prevent, performed by the referee instead of the player.

The rule only means something if it is uncomfortable at least sometimes. A standard you would relax when the outcome embarrasses you was never a standard.

What each layer proves, and what none of them prove, is set out on the trust page.

Every figure above is a count from the record on August 29, 2026, and the archive is exportable if you want to check the arithmetic yourself.

Also here: People who write predictions get scored. People who say them don't. · What makes a prediction testable · What a 96% accuracy score actually measures · How a track record lies without a single false statement · How to check whether someone actually predicted what they say they did · Sometimes the honest verdict is that nobody can tell · Everyone predicts in January and finds out in December