Note

What a 96% accuracy score actually measures

One tracker has Gary Marcus at 96%. Another analysis puts him near 52%. Neither is lying. They counted different things.

A tracker called Forecast Fools has Gary Marcus at 19 predictions tracked, 96% accuracy. A separate analysis of the same man’s claims puts him at roughly 52% supported and 8% contradicted.

Neither is lying. They counted different things.

That is the whole problem with accuracy percentages, and it is not a problem you can fix by counting more carefully. A percentage is a fraction. Somebody chose the denominator, and the person who chooses the denominator decides the answer.

Four decisions hidden in every score

Which claims counted as predictions. Marcus writes constantly. Some of it is dated and falsifiable, some is a mood about where the field is going. Whoever built the tracker drew that line, and no two people draw it the same way.

Which ones were ripe. A tracker can only score what has resolved. If the resolved ones skew toward a certain kind of claim — the short, concrete, near-term ones — the score measures that subset and calls it the person.

What counted as correct. “Pure LLMs will still hallucinate” is obviously right if you mean it happens at all, and arguable if you mean at a rate that matters. Someone decided, usually after seeing what happened.

Who chose the list. The step nobody mentions. A track record assembled by an admirer and one assembled by a critic will differ before a single claim is graded.

Why this registry refuses to publish a score

Not modesty — it is a rule, and it is in the methodology. Profiles here show counts: how many records, how many open, how many settled. Never a percentage.

A count is a fact about the archive. A percentage is a claim about the person, and it inherits every one of the four decisions above. We would be making those decisions and presenting the result as arithmetic.

The second reason is structural. Of the 101 records here, 51 carry a deadline and 50 never will. Any percentage would have to silently drop half the archive — the half where somebody said something forceful with nothing checkable in it, which is often the most revealing half.

What to ask instead

When you meet a track record, the useful questions are not about the number. They are: who picked which claims went in, was that choice made before or after the outcomes were known, and is each entry linked to the original wording so you can check it yourself.

If the answer to the middle one is “after”, the number is a curated selection wearing the costume of a measurement — which is worth a note of its own.

Every figure above is a count from the record on August 29, 2026, and the archive is exportable if you want to check the arithmetic yourself.

Also here: People who write predictions get scored. People who say them don't. · What makes a prediction testable · The goalposts problem · How a track record lies without a single false statement · How to check whether someone actually predicted what they say they did · Sometimes the honest verdict is that nobody can tell · Everyone predicts in January and finds out in December