Base Rate Fallacy

People often underweight general statistical information, the base rate, when a specific, vivid, or representative-seeming case is available, even when that case provides little real evidence and the base rate is strong, though this is closer to a systematic underweighting than a wholesale ignoring of statistics.

4 min read

What Is It?

The base rate fallacy, also called base rate neglect, the more precise term the research literature increasingly favors, describes the tendency to underweight prior statistical probability in favor of specific, individuating information about a particular case. It was demonstrated experimentally by Daniel Kahneman and Amos Tversky in their 1973 paper “On the Psychology of Prediction,” most famously in a study where participants were given a brief personality sketch and asked to judge whether the person described was more likely an engineer or a lawyer; one group was told the sketch was drawn from a pool of 70 engineers and 30 lawyers, another group was told the reverse, 30 engineers and 70 lawyers. Participants’ predictions tracked how representative the sketch seemed of an engineer or a lawyer stereotype far more than they tracked the stated base rate, changing the base rate from 70/30 to 30/70 had strikingly little effect on people’s judgments. This shouldn’t be read as a claim that the sketches were objectively uninformative in some absolute sense; the point is narrower, that the stated statistical context did comparatively little work relative to how strongly the sketch matched a stereotype.

It’s worth being direct about something often left out of the popular version of this finding: the strength of base-rate neglect is highly sensitive to how a problem is presented. People aren’t uniformly blind to base rates, later research has found substantially better use of them under different task formats, when the relevance of the base rate is made clearer, and among people with real domain expertise using base rates adaptively in real-world judgments. The durable finding is that base rates are often underweighted under certain conditions, not that people generally ignore strong statistical evidence whenever a vivid story is available.

The fallacy isn’t a claim that specific case information should always be ignored in favor of the base rate. A base rate itself can be the wrong one to apply, built from a different population than the specific case genuinely belongs to, in which case leaning on it can mislead just as badly as ignoring it does. The actual discipline is checking whether the base rate genuinely applies to the case at hand, then weighing genuinely informative case-specific evidence against it, rather than either defaulting to the base rate or discarding it the moment a vivid story is available.

Why Does It Matter?

The clearest organizational version of this involves a genuine reference class: a hiring manager who believes an unstructured interview has identified an exceptional candidate, while underweighting the historical success rate of similarly evaluated hires in that same role. That’s meaningfully different from simply distrusting a vendor’s own track record, which is longitudinal, case-specific evidence about that particular vendor, not a population base rate; a genuine base rate there would be something like the success rate across comparable vendors or comparable implementations generally, a different reference class than the vendor’s own history. Risk assessment can involve the same error, underweighting a stated base rate, “this failure mode occurs in twenty percent of comparable cases, this one in one percent,” in favor of specific details that make the rarer mode feel more vivid or plausible, which is base-rate neglect specifically when a real statistical prior exists and gets underweighted, and a related but distinct effect, the availability heuristic, when the issue is really about which failure mode comes easily to mind rather than about a stated probability being discounted.

What Changes Once You See It?

You start explicitly asking what the base rate is for a decision before getting pulled in by a specific, compelling case, a typical success rate for this kind of hire, this kind of vendor, this kind of bet, and treating that base rate as a real input rather than background noise the specific story is expected to override. You get more skeptical of a persuasive individual case that doesn’t engage with how it compares to the broader pattern. You also become more careful to check that a base rate you’re relying on is actually drawn from a comparable population to the specific case at hand, rather than simply reaching for whatever base rate is available.

Common Misunderstandings

  • It isn’t a claim that specific case information is always worthless or should be ignored. Genuinely informative case-specific evidence should shift a judgment away from the base rate; the fallacy is about weak or unrepresentative case information overriding a base rate for no good evidential reason, not about all case information being suspect.
  • It isn’t solved by mechanically always favoring the base rate over the specific case. A base rate built from a different or poorly matched population can mislead as badly as ignoring the base rate entirely, so the real discipline is checking whether the base rate actually applies before leaning on it.
  • It doesn’t mean vivid, story-based information is inherently untrustworthy. The problem is the weight it receives relative to its actual evidential strength, not the fact that it’s vivid or memorable; a vivid story can also happen to be strong evidence.
  • It isn’t limited to formal statistical settings. The same dynamic plays out informally whenever a compelling individual anecdote is allowed to quietly substitute for a look at the broader pattern, which happens constantly in everyday organizational decisions that never involve an explicit probability estimate at all.

Diagnostic Question

What does the relevant reference class tell us before we consider this specific story, and how diagnostic is the case-specific evidence that would actually justify moving away from it?

Explore Further

Field Notes

  • None yet.

Related Field Guide

Origin

Demonstrated by Daniel Kahneman and Amos Tversky in “On the Psychology of Prediction,” Psychological Review (1973), which included the widely cited “lawyers and engineers” experiment; later research has substantially qualified how consistently the effect appears across different task formats and levels of expertise.

Know someone who’d enjoy this?