Eight pairs. Each pair has one finding that survived a high-powered replication and one that collapsed. Pick the survivor.
Every outcome here is published — Many Labs, the registered replication reports, the Reproducibility Project. You're being scored against the record, not against opinion.
About ninety seconds. ← and → work.
1 / 8which one replicated?
Result
Here's how you did.
your eight callsgreen = right · amber = wrong
—
Accuracy
How often you picked the finding that actually held up.
—
Fame trap
How often you backed the more famous finding and it turned out to be the one that collapsed.
What you're being compared against
39%
of 100 psychology studies replicated successfully in the Reproducibility Project, and the replication effects averaged about half the original size.
62%
of 21 social science experiments from Nature and Science replicated, in the Social Sciences Replication Project.
61%
of 18 experimental economics studies replicated in a parallel project.
50/193
experiments reproduced successfully in the Reproducibility Project: Cancer Biology, across 23 of 53 papers.
This is a real skill, and it's been measured. Before the Social Sciences Replication Project ran, researchers bet on the outcomes in prediction markets. The markets were highly accurate at calling which findings would survive. People can see this coming — which is exactly why it's worth knowing where you sit.
Your misses
Sources and how to extend this. Outcomes come from Many Labs 1 (Klein et al., 36 samples, 6,344 participants), the registered replication reports on ego depletion (Hagger et al., 23 labs) and facial feedback (Wagenmakers et al., 17 labs), and the aggregate figures from the Reproducibility Project: Psychology, the Social Sciences Replication Project (Camerer et al., 2018), the Experimental Economics Replication Project, and the Reproducibility Project: Cancer Biology.
Eight pairs is demo length. The Reproducibility Project posted outcomes for all 100 studies on OSF — that's your path to a few hundred pairs without asking anyone's permission. Two gaps to close before launch: there are no attention-check pairs in this build, and the crowd field on each pair is my estimate rather than measured, so the fame-trap number is directional until you've logged real sessions.