LESSON
1.03

The three ways we learn, and the one that hooks you

We learn by association, by consequences, and by watching others. The catch is hiding in the consequences: a reward you cannot predict is the hardest habit to quit.

WRITTEN BY
Mike Popesku
PUBLISHED
September 6, 2026

What the science says

Consensus

There are three well-established routes by which experience changes behaviour. The first is association. Pavlov's dogs learned that a bell predicted food and began to salivate at the mere sound of the bell; Watson and Rayner (1920) showed the same machinery could attach fear to a once-neutral object (the "Little Albert" study, which we will come back to, because it should never have been run). The modern correction is the important part. Conditioning is not simply about two things happening close together in time. Rescorla (1988) showed it is about prediction: a cue conditions only when it genuinely forecasts the outcome. A bell that rings just as often but carries no information teaches nothing. In other words, learning is the brain building a model of what predicts what.

The second route is consequences. Thorndike's law of effect and Skinner's operant work established the basic rule: an action followed by reward grows more likely, one followed by punishment less so. The detail that matters most in practice is the schedule on which rewards arrive (Ferster & Skinner, 1957), simply the rule for how often, and how predictably, a reward shows up. Reward every single time and behaviour is brisk but quick to stop when the rewards end. Reward on a variable-ratio schedule, after an unpredictable number of attempts, and you get the highest, steadiest rate of all, and behaviour that is remarkably hard to extinguish, meaning slow to fade once the rewards stop. That is not an accident of the lab. It is the exact schedule a slot machine runs, and the one behind pull-to-refresh feeds and loot boxes.

The third route is observation. Bandura, Ross and Ross (1961) let children watch an adult attack an inflatable "Bobo" doll, then left the children alone with it. Many copied the aggression in detail, despite never being rewarded for it. We learn by watching models, especially models who look like us or who get rewarded.

Underneath all three routes sits one neural theme worth slowing down for. Schultz, Dayan and Montague (1997) found that dopamine neurons do not simply fire for reward. They fire for surprise. A dopamine burst marks the gap between the reward you expected and the reward you actually got. Something good you did not see coming fires them hard. Once a cue reliably predicts that reward, the burst shifts back to the cue, and the reward itself stops producing one. And when an expected reward fails to arrive, they dip below baseline, a small signal of disappointment. The system is not tracking pleasure; it is tracking prediction error, the difference between what you expected and what happened. That is exactly why an unpredictable reward keeps you hooked: it never stops generating surprise, so it never stops teaching.

Controversies

Three caveats matter. The "Little Albert" study is a landmark and an ethical disgrace: a single infant, frightened deliberately, never de-conditioned, with later disputes about who he even was. We cite it for historical completeness, not as a suggested method. More broadly, the mid-century belief that conditioning explains essentially all learning overreached; language acquisition and one-trial learning do not fit it, and biology constrains it (the Garcia effect: a single pairing of an unusual taste with later illness produces a powerful aversion, because animals are prepared to link food with sickness and not, say, with a sound). And the popular phrase "a dopamine hit" misreads the science: dopamine tracks prediction error and wanting, not pleasure itself.

Limitations

The cleanest schedule findings come from animals pressing levers, and humans can override a schedule with language, rules, and conscious goals. Effects also vary with how prepared we are to make a given association. The principles are robust; their exact reach into messy human life is an extrapolation, not a guarantee.

Open questions

The live frontier is the line between fast, habit-like "model-free" learning and slower, goal-directed "model-based" learning, and how the prediction-error system interacts with deliberate thought.

So what

Behaviour can be trained, yours and other people's, and the levers are cue, consequence, schedule, and example. The single most powerful and most abusable of these is the schedule.

For companies

Reward schedules are engagement design, whether you intend them or not. Predictable rewards (a point for every purchase) build brisk but fragile habits. Unpredictable rewards (a surprise upgrade, a mystery deal, the slot-machine logic of a feed that sometimes delivers something great) build the stickiest behaviour of all, and that power is strong enough to raise a genuine ethics question, which the applications layer returns to. There is also a gentler lesson: because learning is prediction, a brand that consistently predicts a good outcome becomes a trusted cue, and consistency is what builds that link. And people copy models, so showing customers like them using the product teaches more than describing it.

For political parties

Activism is operant behaviour: it is sustained by reinforcement, and intermittent, unpredictable wins and recognition keep people going longer than a steady, expected drip. Observation is just as strong a lever: people are far more likely to donate, attend, or vote when they see people like themselves already doing it, which is why visible social proof beats abstract appeals.

For government

Incentives change behaviour, but the schedule and the source matter. One-off rewards extinguish fast once they stop. Punishment works narrowly and brings side effects, so it is a blunt tool. The underused lever is modelling: showing that most people already comply (paying tax on time, for example) teaches the behaviour by observation more cheaply than a threat does.

How to use this

Stop thinking "how do I motivate this?" and start thinking in four parts: the cue that predicts the behaviour, the consequence that follows it, the schedule that consequence arrives on, and the example people can see. Unpredictable reward builds the most durable behaviour, so use it deliberately and ethically. And remember the brain learns from surprise, so the unexpected reward, not the routine one, is what teaches.

Why you can't stop

Case studies

References

  1. Watson & Rayner (1920), J. Experimental Psychology. DOI 10.1037/h0069608
  2. Bandura, Ross & Ross (1961), J. Abnormal & Social Psychology. DOI 10.1037/h0045925
  3. Rescorla (1988), American Psychologist. DOI 10.1037/0003-066X.43.3.151
  4. Schultz, Dayan & Montague (1997), Science. DOI 10.1126/science.275.5306.1593
  5. Pavlov (1927), Conditioned Reflexes (book)
  6. Thorndike (1911), Animal Intelligence (book)
  7. Skinner (1953), Science and Human Behavior (book)
  8. Ferster & Skinner (1957), Schedules of Reinforcement (book)
  9. Bandura (1977), Social Learning Theory (book)