LESSON
2.06

Cognitive Abilities - What intelligence measures, and what it misses

Of all psychology's ideas, intelligence is among the best measured and the most abused. There is a real, sturdy finding underneath it. There is also a century of people stretching that finding into things it never showed.

WRITTEN BY
Mike Popesku
PUBLISHED
September 6, 2026

What the science says

Consensus

Start with the one fact everything else is built on. If you give people a pile of different mental tasks, vocabulary, number patterns, mentally rotating a shape, remembering a sequence, the scores all correlate positively. Someone good at one tends to be good at the others, even when the tasks look unrelated. Charles Spearman named this the positive manifold and proposed that a single general factor sat underneath it, which he called g (Spearman, 1904). The positive manifold is about as solid as findings in psychology get, replicated across more than a century, countless test batteries, and every population studied.

On top of that base sits a tidy structure. Raymond Cattell split ability into fluid intelligence, reasoning your way through a novel problem, and crystallised intelligence, the knowledge you have accumulated (Cattell, 1963). John Carroll's survey of more than 460 datasets pulled the whole field into a three-layer map, since merged into what is now the consensus model: g at the top, around ten broad abilities beneath it (fluid reasoning, crystallised knowledge, memory, processing speed, and so on), and dozens of narrow skills below those (Carroll, 1993). One general factor, several broad ones, many specific ones.

These tests do real work. Scores are reliable and strikingly stable across a lifetime, and they predict outcomes that matter: childhood ability tracks national exam results at correlations around .5 to .8 (Deary et al., 2007), and across the lifespan measured intelligence is linked to education, occupation, health, and even how long people live (Plomin and Deary, 2015). Not perfectly, not for any one individual, but reliably on average. This is the genuine, hard-won core. Now the trouble.

Controversies

The first crack runs through the most-quoted statistic in the field. For decades the headline was that cognitive ability is the single best predictor of job performance, a correlation around .5, from a landmark meta-analysis (Schmidt and Hunter, 1998). In 2022 a team rechecked the statistical corrections that produced those numbers and found them systematically overdone: a standard adjustment for "range restriction" had been inflating the validity of cognitive tests for years (Sackett et al., 2022). Corrected, the estimates fell by .10 to .20, and cognitive ability lost its top spot to the structured interview. It still predicts performance. It is simply not the juggernaut a generation of textbooks claimed, and the episode is a clean example of a field correcting its own celebrated number.

The second crack is the Flynn effect. Across the 20th century, raw IQ scores rose by roughly three points a decade in country after country (Flynn, 1987). Genes do not change that fast, so something environmental, better nutrition, schooling, more abstract daily demands, was clearly lifting measured intelligence. And the trend is not one-way: in several wealthy countries scores have stalled or slipped in recent decades, and a study tracking Norwegian brothers showed the rise and the reversal happening within families, which pins both on environment rather than genes or changing demographics (Bratsberg and Rogeberg, 2018). Whatever the tests catch, it is movable.

The third crack is the one people fight about: heritability. Twin and genomic studies agree that intelligence is substantially heritable, and, oddly, that heritability climbs with age, from around 20% in early childhood to perhaps 80% in later adulthood (Plomin and Deary, 2015). This is where misreading does real damage, so it is worth being precise. Heritable does not mean fixed: the Flynn effect moved whole populations within a generation, and adopting children from poorer into better-off homes raises their scores by 12 to 18 points (Nisbett et al., 2012). Heritability is also a within-group statistic that depends on the environment it is measured in, it is lower among children raised in poverty, where bad environments cap potential before genes get to express it (Nisbett et al., 2012). A trait can be highly heritable within groups and still be pushed around massively by circumstances.

That precision matters most for the fourth and most charged crack: group differences and bias. Average differences between groups on these tests are real and measured. Two things are then usually overclaimed. First, that the tests must be biased: in the technical sense they largely are not, a well-built test predicts later outcomes about equally well across groups, which is what "predictive bias" means. Second, and far more important, that heritability explains the gaps: it does not. Heritability within a group says nothing about what causes the average difference between groups, the gaps have narrowed over time as circumstances changed (the Black-White gap in the United States shrank by about a third of a standard deviation; Nisbett et al., 2012), and the causes remain genuinely unresolved. Given that this is the corner of the field with the longest history of pseudoscientific misuse, the honest position is narrow: real differences in scores, contested and unsettled causes, and no warrant for the hereditarian leap.

Two smaller disputes round it out. There is an old argument about what g even is, a single underlying mental engine, or a statistical shadow cast by many overlapping skills that could, in principle, grow up reinforcing one another (Nisbett et al., 2012). And the popular alternatives that promise to dethrone IQ, Gardner's multiple intelligences and the idea of visual or auditory "learning styles", have not held up to measurement; g keeps reappearing whenever the alternatives are tested properly (Gardner, 1983).

Limitations

Whatever its strengths, the test captures a real but narrow slice of the mind, and leaves out a great deal we also call being smart: wisdom, creativity, practical judgement, character. Crystallised measures carry cultural and educational fingerprints. The scores predict averages, not the person in front of you. And a test result is co-produced by motivation, anxiety, sleep, and practice, so it is a sample of behaviour on a particular day, not a readout of fixed capacity (L2-08, L0).

Open questions

What is g, physically, in the brain? Can intelligence be raised in a way that lasts, given that most training gains fade once the training stops? What actually causes the average differences between groups? And is the general factor a real entity or a useful summary of skills that travel together?

So what

The usable core: measured ability is a real, narrow, partly movable signal that predicts averages, and almost every practical mistake comes from treating it as fixed, broad, or a measure of human worth.

For companies

Cognitive ability genuinely predicts who will do well, which is why it is tempting in hiring, but the recent correction is a direct warning against leaning on it too hard: its real-world validity is lower than the famous numbers claimed, and the structured interview now outranks it (Sackett et al., 2022). The sound approach is to combine signals, a work sample, a structured interview, and an ability measure, rather than chasing "the smartest candidate" on a single test. There is also a practical fairness and legal dimension: cognitive tests show average group differences, so relying on them heavily narrows your hiring diversity and invites challenge, which is part of why the validity-versus-diversity trade-off is worth designing around rather than ignoring.

For political parties

Intelligence is one of the easiest pieces of science to weaponise, as a marker of innate worth, a justification for who deserves what, or a cudgel in arguments about groups. The research does not support those uses. The thing is narrow, it moves with environment, and its most explosive claims (that gaps are genetic and fixed) are exactly the ones the evidence does not establish. Meritocratic rhetoric in particular often rests on a fixed-IQ myth that the Flynn effect quietly refutes: if intelligence can rise three points a decade on better conditions, it was never the pure, unearned birthright the story needs it to be.

For government

The same evidence that makes intelligence sound deterministic actually argues for investment. Environments move it: nutrition, schooling, and pulling children out of poverty all raise measured ability (Flynn, 1987; Nisbett et al., 2012), which gives early-childhood and education spending a real cognitive payoff. The flip side is a caution against sorting and tracking children early on a narrow, still-moving measure, and against reading the genuine links between ability and health or lifespan as fate rather than as one more reason to fix the environments that shape both.

How to use this

Three habits. First, treat a test score as one real but narrow signal, a sample of certain skills on a given day, never a verdict on a person's worth or potential. Second, use more than one signal for any decision that matters, since the single-number approach is both weaker and less fair than it looks. Third, hold the line against the two big overreaches: that intelligence is fixed (it visibly moves), and that heritability explains gaps between groups (it does not).

Find the hidden factor

Four very different mental tests, each tapping a different ability. Twelve people sat all four. Pick one test for each axis to see how those two sets of scores line up, then try another pairing, and another.

Across
Up
Correlation: --

Each dot is one person. Uphill means people who score high on one test tend to score high on the other.

Pairs explored: 0 of 6
What you just found

One honest caveat: a single factor that summarises all that overlap is real and useful, but it does not by itself prove there is one inner "engine" driving every skill. Whether g is a thing or a shadow cast by many overlapping abilities is still argued. Either way, the pattern you just traced is about as solid as findings in psychology get.

The universal positive correlation across mental tests is the positive manifold; Spearman (1904) proposed the general factor g to explain it. Illustrative data, but it behaves like real test batteries, where every pair of well-built cognitive tests intercorrelates positively.

Case studies

References