Some of the most famous findings in behavioural science fell apart when they were redone. That isn't a scandal, it's quality control, and it tells you exactly how much to trust a single study.
In 2015 the Open Science Collaboration published the result that named the problem. Independent teams repeated 100 studies from three top psychology journals. About 97% of the originals had reported a clear positive result; only around 36% of the repeats did, and even the effects that held up came in at roughly half their original size. A few years later, Camerer and colleagues (2018) redid 21 social-science experiments from Nature and Science and found 13 replicated, again at about half strength. So the picture is not "nothing is real." It is "most of the careful, top-tier work holds, and a meaningful slice does not."
Some specific casualties are worth naming, because they were famous. Ego depletion, the idea that willpower runs down like a fuel tank, was a textbook effect with thousands of citations (Baumeister et al., 1998). When 23 labs ran the same preregistered test on roughly 2,000 people, the effect came out essentially zero (Hagger et al., 2016). Power posing fared a little better and a lot worse: a larger study found that standing in a confident pose made people feel more powerful but did not move their hormones or their risk-taking (Ranehill et al., 2015), and the original's first author later said publicly that she no longer believed the effect was real. Social priming results, like the classic finding that words about old age made people walk more slowly, failed to reproduce, with the experimenters' own expectations implicated (Doyen et al., 2012). The causes turned out to be ordinary: small samples, flexible analysis choices that quietly manufacture significance (Simmons, Nelson & Simonsohn, 2011), and a publishing system that prizes the surprising over the solid, exactly as Ioannidis (2005) had warned.
How bad is it, really? That part is genuinely disputed. Gilbert and colleagues (2016) reanalysed the Open Science Collaboration data and argued that once you account for low statistical power and the fact that many repeats used different populations, the results are consistent with reproducibility being quite high. The original team pushed back, and the exchange sharpened how the field now defines a "successful" replication. So the headline number moves depending on how you count. What almost nobody disputes is the direction: a real chunk of the published literature is fragile, and the methods that produced it were too loose.
A failed replication is not a verdict on its own. It can mean the original was a false positive, or that the new study differed in population, context, or sample size. This is why single replications rarely settle anything and multi-lab preregistered efforts carry the weight. It is also why the crisis is best understood as concentrated in specific areas, mainly social and cognitive psychology and parts of biomedicine, rather than a blanket indictment of "science."
We still don't know the true replicability rate for most subfields, or how much of the failure reflects false originals versus effects that are real but fussy about context. And the big practical question is open: whether the reforms now spreading, preregistration, registered reports, and bigger samples, are actually raising the hit rate. The early signs are encouraging, but the measurement is ongoing.
This is the discipline that earns everything else in this series its credibility. If you are going to act on a behavioural finding, you need a way to tell the durable ones from the fashionable ones. The replication crisis hands you that filter.
Be wary of strategy built on a single exciting study. The findings that made the best conference talks (power posing, subtle priming, willpower as a fuel tank) are among those that failed to hold up. Before you spend budget on a tactic, ask three questions: has it replicated, how big is the effect when it does, and was the test preregistered? A flashy result from one small study with a huge effect is a warning sign, not a green light.
Campaign folklore is full of one-study tricks: a subliminal cue, a single clever message that supposedly swung a race. Most of these do not survive contact with a large field experiment. Trust effects that show up repeatedly, at scale, in real elections, and discount the cute lab result that a vendor is selling as a silver bullet.
Policy is where fragile findings do real damage, because they get deployed at scale. Lean on replicated, field-tested evidence, expect effects to shrink when they leave the lab, and build evaluation into the rollout so a weak intervention gets caught early rather than scaled blindly.
Treat every single study as provisional. Weight evidence by whether it has replicated, by the size and stability of the effect rather than mere statistical significance, and by whether the design was preregistered. The rule of thumb that will rarely fail you: prefer boring and robust to exciting and fragile.