The first version was too easy, and the reason was a design flaw: the failed findings were mostly odd priming studies and the survivors were mostly classic judgment effects. You could score well without knowing anything, just by pattern-matching on vibe.
So every pair here is matched — same field, usually the same journal and roughly the same years. Two cooperation experiments from Nature. Two classroom interventions from Science. Two embodiment studies from 2010. One of each pair survived a high-powered replication. One didn't.
Fifteen pairs. ← and → work. Take your time — rushing is detected.
Before the Social Sciences Replication Project ran, researchers bet on the outcomes in prediction markets, and the markets called it well — expert forecasters correctly predicted replicability for close to three quarters of the studies. That result is why a test like this seemed worth building.
But SCORE, finishing in 2026 on a far larger and more varied sample, found the opposite: no single indicator of trustworthiness stood out, expert and machine predictions related only modestly to actual replicability, and AI models trained to imitate accurate human forecasters still failed to pick the winners. The one factor that did correlate strongly with reproducibility was mundane — whether the data and code had been made available.
Which means your score above is genuinely ambiguous. On a curated set of famous studies, forecasting works. Across the literature as a whole, it may not. Both things appear to be true, and nobody has fully reconciled them.