Facts / replication
The claim
In 2015, 100 psychology studies were re-run; only 36% replicated.
AI confidence
85%
that the claim as stated is accurate.
- •The numbers match the paper: 100 replications were reported in 2015, and 36% of replications reached significance.
- •"Replicated" is one of five criteria the authors used; the others give 39%, 47% and 68%, so "only 36% replicated" is accurate but picks the strictest common reading.
- •The replications were run between 2011 and 2015, and 36% is of the 97 originals that were significant, not of all 100; both are small imprecisions.
- •The claim is about what was observed, not about how many original findings are false, and should not be read as the latter.
Assessed by Claude, 2026-10-04.
Your confidence
How confident are you that the claim is true?
Give a probability from 0 to 100. One answer per browser.
The evidence
Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science 349(6251), aac4716.
The Reproducibility Project: Psychology repeated 100 experimental and correlational studies published in 2008 in three journals: Psychological Science, Journal of Personality and Social Psychology, and Journal of Experimental Psychology: Learning, Memory, and Cognition.
The replications used high-powered designs and the original materials when available, and original authors were invited to review the protocols in advance.
About 270 researchers took part, coordinated by the Center for Open Science, and the results were published in Science in August 2015.
| Original studies with a significant result (p < .05) | 97% |
| Replications with a significant result | 36% (35 of 97) |
| Mean effect size, original studies | r = 0.403 |
| Mean effect size, replications | r = 0.197 (about half) |
| Original effect inside the replication 95% CI | 47% |
| Rated as replicated by the replication teams | 39% |
| Significant after combining original and replication data | 68% |
What "replicated" meant
The paper says there is no single standard for replication success and reports five indicators.
The 36% figure is the significance criterion: the replication found a statistically significant effect in the same direction.
By the other criteria the rate was 39% (subjective rating), 47% (original effect inside the replication confidence interval) and 68% (original and replication data pooled, which assumes the original estimate is unbiased).
A failed replication does not prove the original effect is zero.
It means the new evidence was weaker than the original evidence suggested.
The main critique
Gilbert, King, Pettigrew and Wilson (Science, 2016) argued the paper contains three statistical errors: it ignored error from differences between original and replication protocols, it relied on a single replication attempt per study, and some protocols were not endorsed by the original authors.
They reported that replications with endorsed protocols succeeded far more often than unendorsed ones, and concluded the data are consistent with fairly high reproducibility.
The authors’ reply
The project authors (Anderson et al., Science, 2016) answered that the critique rests on statistical misconceptions and causal readings of correlational data, since endorsement was not randomly assigned.
They argued that both very optimistic and very pessimistic conclusions are possible from these data, and that neither is yet warranted.
Both sides agree on the counts.
They disagree on what the counts say about the true rate of reproducible findings.
Sources
- Open Science Collaboration (2015), Science 349(6251), aac4716https://doi.org/10.1126/science.aac4716
- Reproducibility Project: Psychology, OSF project (data, protocols, analyses)https://osf.io/ezcuj/
- Gilbert, King, Pettigrew, Wilson (2016), Comment, Science 351(6277), 1037https://doi.org/10.1126/science.aad7243
- Anderson et al. (2016), Response to Comment, Science 351(6277), 1037https://doi.org/10.1126/science.aad9163
Related: theburningmind on YouTube · hbar.systems