LESSON
5.13

Nudge and Choice Architecture - Nudge, and what's now embarrassing

Nudging was the behavioural idea that conquered government: small, cheap tweaks to how choices are presented, steering behaviour while leaving people formally free. Then the evidence caught up. The honest position now is that a few nudges work powerfully, most work barely at all, and the field oversold the difference for a decade.

WRITTEN BY
Mike Popesku
PUBLISHED
September 6, 2026

What the science says

Consensus

Nudge crystallised a genuinely useful idea (Thaler and Sunstein, 2008). A choice architect, anyone who designs how options are presented, is never neutral, so you may as well arrange the defaults, the framing, the order, and the salience to steer people toward better outcomes while leaving every option formally available. The approach was cheap, politically palatable (it preserved freedom of choice), and it spread fast: the UK Behavioural Insights Team formed in 2010, and within a few years dozens of governments and the OECD had their own units (Halpern, 2015).

Several of the flagship results were real. Automatic enrolment turned pension participation from a minority to a near-universal default (Madrian and Shea, 2001, treated in L5-09), and countries with an opt-out organ-donation default show dramatically higher registration than opt-in ones (Johnson and Goldstein, 2003). Tax authorities ran letter experiments that reliably lifted payment. The common thread of the durable wins, it turns out, is that they were mostly defaults and simplifications, not persuasive messages.

Controversies

The reckoning began around 2020 and it is the heart of the story. A prominent meta-analysis pooled hundreds of choice-architecture studies and reported a solid average effect (Mertens et al., 2022). A reanalysis by Maier and colleagues then corrected the same data for publication bias, the tendency for only significant results to get published, and found that the average nudge effect was statistically indistinguishable from zero (Maier et al., 2022). The exchange became the field's public reckoning with its own evidence base.

The most credible number came from inside the movement. DellaVigna and Linos analysed 126 randomised trials run by two real government nudge units on roughly 23 million people, the closest thing to an unbiased sample of what nudging actually does when a unit deploys it at scale. The average effect was about 1.4 percentage points on a 17-point baseline, around an 8% relative improvement, against an average of roughly a third in the published academic literature (DellaVigna and Linos, 2022). That several-fold gap is the publication-bias and selective-reporting story made concrete: the nudges that get written up are not the nudges you will get.

None of this means nudging fails. It means the field oversold a class of interventions whose average effect is small, and it consistently overweighted the weak members of the family (social-norm signs, pop-up prompts, exhortation) relative to the strong ones (defaults, simplification, timely reminders). The genuinely embarrassing part is the second claim that grew up alongside the movement: that low-cost behavioural tweaks could stand in for structural policy on health, poverty, or climate. On the evidence, they cannot.

Limitations

The published effect sizes are inflated by publication bias and by the gap between a tight academic study and an at-scale field deployment. Effects are highly heterogeneous, defaults dwarf prompts, so any "average nudge effect" is close to meaningless as a planning number. There are ethical limits too: the same friction that can be removed to help people can be added to trap them (sludge), and the paternalism question, who decides which way to steer, never fully goes away. And the evidence is culturally narrow: the famous organ-donation default gap reflects registration systems and interacts with the surrounding medical infrastructure, so a default that works in one country is a hypothesis, not a guarantee, in another.

Open questions

Which nudges scale, and can we predict in advance which will work rather than discovering it in a trial? Where exactly is the line between a legitimate nudge and manipulative sludge or a dark pattern? And how much of the apparent success of the best nudges is the default effect doing nearly all the work?

So what

The usable core: build the few nudges that genuinely work, discount the published effect sizes several-fold when you forecast, test your own version, and never treat a nudge as a substitute for structural policy. What follows is a triage of the family, strongest to weakest, with the honest effect size and the failure condition for each.

The ethics run underneath all of it. Choice architecture is unavoidable, someone always arranges the options, so the real question is whose interest the arrangement serves. A default or a friction that helps the person choosing is a nudge; the identical technique turned to trap or exhaust them (sludge, the cancel flow that takes twenty clicks, the pre-ticked box that costs them money) is the same tool used against them. The honest test is the one running through this series: would the design survive the person seeing exactly how and why it was built?

The playbook: which nudges to build, and which to distrust

Defaults, the one big, reliable lever. This is where most of nudging's real power lives. Set the outcome that is genuinely good for the person as the default and let inertia do the work: opt-out pension enrolment, default renewable energy tariffs, automatic registration. Honest effect: large where it applies, many times bigger than the prompt-style nudges. Failure condition and ethics: a default is only legitimate when the defaulted option is genuinely in the person's interest, and defaults interact with the surrounding system, an opt-out organ-donation register raises registration but does not by itself raise actual donations without the transplant infrastructure behind it, so do not confuse the default with the outcome. The mechanics and the opt-in/opt-out evidence are treated in depth in L5-16.

Friction reduction and simplification, the cheapest real win. Cut steps, pre-fill fields, shorten and clarify forms, remove the small hassles between a person and the thing they already want to do. Honest effect: reliable and often the best return on effort, because you are removing a barrier rather than trying to manufacture motivation. Failure condition: it only helps when friction was the actual obstacle; simplifying a form does nothing if the person never wanted the outcome.

Timely reminders, real but modest, and only for the already-motivated. A reminder delivered at the moment of action lifts follow-through for people who intend to act but forget. Honest effect: genuine but small, and this is where the DellaVigna and Linos ~8% benchmark should set your expectations, not the literature's 33%. Failure condition: a reminder fixes a memory or timing gap, not a motivation gap (in COM-B terms, L5-14, it addresses capability and opportunity, not motivation), so reminding someone to do a thing they do not want to do achieves nothing.

Social-norm signs, pop-up prompts, and exhortation, mostly distrust these. This is the weak end of the family that the field overweighted. Honest effect: small on average, frequently indistinguishable from zero once publication bias is stripped out (Maier et al., 2022). Failure condition: worse than nothing when the norm cited is unfavourable, since a sign announcing that the bad behaviour is common increases it (the petrified-forest backfire, L5-12 and L5-18). If your behavioural strategy rests on signage and prompts, you are relying on the part of the toolkit least likely to work.

The meta-directive: forecast honestly and test. When you plan against a nudge, expect several times less than the published effect size and use the at-scale nudge-unit figures as your prior (DellaVigna and Linos, 2022). Run your own randomised trial rather than trusting a headline result, because the nudge that was written up is a biased sample. And hold the hard line the reckoning established: a nudge is a complement to good policy, never a substitute for it. If the problem is structural, budget accordingly and use the nudge at the margin.

For companies

Concentrate your behavioural effort on defaults and friction reduction, the parts that actually move numbers, and treat on-site prompts and social-proof widgets as low-value until proven otherwise in your own test. Set sensible, genuinely-good-for-the-customer defaults, strip friction out of the paths customers want to travel, and add sludge audits to find the friction you have accidentally (or deliberately) left in the paths they want to leave, because a punishing cancellation flow is a reputational and increasingly a regulatory liability. Forecast with the honest effect sizes: a proposal promising a large behavioural lift from a norm message or a pop-up is selling the literature's inflated number, not the one you will get.

For political parties and campaigns

The usable nudges in campaigning are the reliable ones: making the desired action easy (simplified registration, clear polling information delivered at the right time) and well-timed reminders to already-motivated supporters, which is the real mechanism behind a lot of effective get-out-the-vote work (L5-01). Be sceptical of consultants promising large persuasion effects from clever message framing alone; the honest expectation is small. And resist the temptation to present cheap behavioural tactics as a substitute for the harder work of organisation and a genuine offer.

For government and public services

This is nudging's home and its cautionary tale. The behavioural-insights units earned their place with defaults, simplification, and timely reminders, and those tools remain genuinely worth deploying across tax, health, and public services. Two disciplines now matter more than the early enthusiasm allowed. First, plan with the real at-scale evidence: the DellaVigna and Linos ~8% relative benchmark, not the academic literature's headline, is the number to budget against, and any nudge should be run as a trial with an honest counterfactual (the units' own best practice). Second, hold the line that nudges complement rather than replace structural policy: a decade of enthusiasm let cheap behavioural tweaks be offered as an answer to problems, from obesity to poverty to emissions, whose real drivers are structural, and the evidence does not support that substitution (this is the same structure-over-exhortation lesson as L5-08). Run sludge audits of citizens' journeys, since the friction a state accidentally builds into benefits or services can quietly exclude the people who need them most. The global caveat is sharp: the flagship default results reflect specific national systems, so effect sizes and even directions should be treated as priors to test locally, not imported.

How to use this

Three habits. Build the strong nudges, distrust the weak ones: defaults, simplification, and timely reminders earn their place, while norm signs and pop-ups mostly do not. Expect several times less than the headline and test your own version, because the published nudge is a biased sample and the at-scale reality is roughly 8%, not a third. And keep nudges in their place, as a complement to structural policy, never a substitute, and audit your own designs for the sludge that turns a helpful tool into a trap.

The incredible shrinking nudge

A nudge is a small, cheap change to how a choice is presented, a reminder, a default setting, a better-worded letter, meant to gently steer what people do. The real question is how much a nudge actually works. Watch one number survive a reality check.

Picture a nudge aimed at a real goal, say getting more people to pay a bill on time. In the studies that got published, nudges like this lifted the behaviour by about a third, on average. The bar shows the size of that effect.
the headline
no effectdoubles the behaviour
Published studies: about a third more

Case studies

References