LESSON
0.04

Every way to measure people lies a little

Surveys, clicks, taste-tests, reaction-time tests: each one has a built-in blind spot. The skill is combining them so their errors don't line up.

WRITTEN BY
Mike Popesku
PUBLISHED
September 6, 2026

What the science says

Consensus

Every way of measuring behaviour has a signature strength and a signature blind spot. Knowing both is what makes the difference from succeeding in your research project or failing.

Self-report (surveys and interviews) is the workhorse because it is cheap and scales. It also carries the most baggage: people forget, people present a flattering version of themselves, and there is a subtler trap called common-method bias. When you measure the supposed cause and the supposed effect in the same questionnaire, the correlation between them is inflated by the shared method alone, before any real relationship is considered (Podsakoff et al., 2003). In other words, a respondent will see a connection between questions within the same questionnaire simply because they are in the same questionnaire. Behavioural traces, the logs, transactions, and records people leave behind, fix the honesty problem but introduce another: they only capture what was recorded, and they say nothing about 'why' and even less about the 'why not' (why they did X instead of Y). When Scharkow (2016) checked people's self-reported internet use against their actual log data, the two diverged substantially. People are poor narrators of their own behaviour.

Then there are the sharper instruments. Cognitive interviewing sits a respondent down and probes how they actually understood a question, which catches broken survey items before they quietly ruin a dataset (Beatty & Willis, 2007). Conjoint analysis forces people to make trade-offs between bundled options, which reveals what they truly weigh far better than asking "how important is price to you?" (Green & Srinivasan, 1978). And implicit measures, above all the Implicit Association Test, use reaction times to try to capture attitudes people will not or cannot put into words (Greenwald, McGhee & Schwartz, 1998). The unifying lesson is old and well-evidenced: no single method is trustworthy on its own, you must triangulate. Combine methods whose errors point in different directions, and believe the result that survives across several of them (Campbell & Fiske, 1959).

Controversies

The loudest dispute is about the IAT. It became enormously popular as a measure of hidden bias, and it clearly captures something. The question is whether it predicts what an individual person will actually do. The careful answer is 'mostly no'. Oswald and colleagues (2013) found the IAT predicted real discriminatory behaviour poorly, and no better than simply asking people their views. Forscher and colleagues (2019), pooling 492 studies and over 87,000 people, found you can nudge someone's implicit score a little, but that shift does not carry through to their behaviour. Defenders, including the test's originators, argue it still tells you something real at the group level. The workable consensus: useful in aggregate, unreliable as a verdict on any one individual.

Limitations

Every method here is a proxy, including the behavioural ones, because a log records the action but not the motive. Conjoint only works if the trade-off task resembles the real decision. Cognitive interviewing is typically slower and it's qualitative, so it takes longer to clean, prepare, and process. And even triangulation (the fix) is not a 100% guarantee: methods often disagree, and when they do you still have to make a judgement rather than read off an answer. The honest default, when self-report and behaviour conflict, is to weight the behaviour. This is, in a nutshell, why we, humans, still haven't figured out an omnicomprehensive surefire methodology that is as strong as physics theorems. A mathematical formula in physics will tell you exactly what the outcome of some physical phenomenon will be, with minimum degree of error. Triangulating survey methodologies will not give you that same degree of certainty. It's just something a market research practioner (and buyers) must come to terms with.

Open questions

We still argue about how best to combine methods, about what the IAT is really measuring, and about how to recover the motive behind trace data without falling back on the unreliable "why did you do that?" question.

So what

If you run your understanding of customers, voters, or citizens off a single instrument, you have bought that instrument's blind spot wholesale. The point of this layer is to stop trusting any one number by itself and start asking what observations would have to agree before you believe them.

For companies

Do not build a strategy on one survey. Put stated data (what people say), revealed data (what they actually do in your logs and tills), and trade-off data (conjoint) side by side, and act on what they agree about. Watch for common-method bias in particular: if you measure "satisfaction" and "intent to repurchase" in the same questionnaire and find they correlate, some of that is the questionnaire talking, not your customers. And treat implicit-bias dashboards with care, because they do not reliably predict what one person will do.

For political parties

A poll is self-report and inherits every weakness of it. After all, we have seen several occasions in recent years where polls were actually totally off compared to the actual voting results. Weight the poll against behaviour (past turnout, registration, donations) and against trade-off-style message testing. Remember: the instrument that might flatter your candidate is not more accurate for flattering them. Clear and honest insights are stronger weapons for your campaign.

For government

Evaluation should combine administrative and behavioural data with surveys, never surveys alone. One cheap, high-return habit: cognitively pre-test survey questions before fielding them, because a misread question produces clean-looking numbers that mean nothing. And be wary of using implicit-bias tests as an individual diagnostic in hiring or training, where the evidence does not support that use. That should include a plethora of administration methods, so to flatten the inherent channel pre-selection bias.

How to use this

Choose methods whose blind spots do not overlap, then look for convergence. The aim is not one perfect measurement, it is several flawed measurements that happen to agree. And when the survey and the behaviour tell different stories, follow the behaviour.

Build your triangulation

Your mission: find out why customers are quietly cancelling. No single tool can tell you. Assemble a kit whose blind spots cover for each other.

To really know, you must uncover

Your toolkit (tap to deploy)

Case studies

References

  1. Campbell & Fiske (1959), Psychological Bulletin. DOI 10.1037/h0046016
  2. Podsakoff et al. (2003), Journal of Applied Psychology. DOI 10.1037/0021-9010.88.5.879
  3. Scharkow (2016), Communication Methods and Measures. DOI 10.1080/19312458.2015.1118446
  4. Beatty & Willis (2007), Public Opinion Quarterly. DOI 10.1093/poq/nfm006
  5. Green & Srinivasan (1978), Journal of Consumer Research. DOI 10.1086/208721
  6. Greenwald, McGhee & Schwartz (1998), JPSP. DOI 10.1037/0022-3514.74.6.1464
  7. Oswald et al. (2013), JPSP. DOI 10.1037/a0032734
  8. Forscher et al. (2019), JPSP. DOI 10.1037/pspa0000160