Simpson’s Paradox

A trend visible in the combined, aggregate numbers can reverse completely once you break the same data down by the groups that actually make it up.

3 min read

What Is It?

Statistician Edward Simpson formally described the pattern in 1951, though earlier statisticians had noticed versions of it: a relationship that holds within every subgroup of a dataset can disappear, or flip direction entirely, once those subgroups are combined into one aggregate figure. It isn’t a data error or a calculation mistake. It happens when the groups being combined differ in their underlying rates and are represented in different proportions across whatever’s being compared, so the aggregate ends up reflecting both actual performance and the underlying mix at once, rather than performance alone. A classic version: a hospital can have a better recovery rate than a competitor for both mild cases and severe cases considered separately, and still show a worse overall recovery rate, because it treats a higher proportion of severe cases, which lowers its blended average even though it’s outperforming at every individual level of severity, a difference in case mix rather than a difference in care.

Why Does It Matter?

Organizations report and act on aggregate metrics constantly, an overall approval rate, an overall customer satisfaction score, an overall conversion rate, without checking whether the underlying groups being blended together are comparable. A company-wide metric can decline even while every individual team’s metric is improving, if the mix of what each team is handling has shifted, more of the harder cases going to the team that’s improving fastest, for instance, dragging the blended average down even as every real underlying trend points up. Leadership reading only the aggregate number draws exactly the wrong conclusion, and can end up correcting a team that’s actually doing better than it was before. This is a distinct risk from ordinary bad data. The numbers can be completely accurate at every level and the aggregate conclusion can still be backwards, because the aggregate mixes together performance and composition in a way the subgroup comparisons do not.

What Changes Once You See It?

You stop trusting an aggregate trend until you’ve checked whether the composition of what’s being aggregated has shifted, since a change in mix can produce the same aggregate movement as a genuine change in underlying performance. You start asking for the subgroup breakdown before reacting to a company-wide or team-wide number moving in a concerning direction, particularly when different segments are known to have very different underlying rates. You also get more careful about comparing two aggregate numbers across time or across teams without confirming the mix behind each one is actually comparable, since an identical-looking metric can be measuring different populations of work. You start separating performance change from mix change as a matter of habit. Before asking why the organization got better or worse, you ask whether the population of work being measured simply became easier or harder.

Common Misunderstandings

  • It isn’t a claim that aggregate metrics are useless. The paradox occurs under a specific, checkable condition, subgroup composition that differs across whatever’s being compared, not a universal warning against summarizing data.
  • It isn’t the same as cherry-picking data to support a conclusion. Simpson’s Paradox can appear in completely honest, correctly calculated numbers, the reversal is a real mathematical property of how the groups combine, not a sign of manipulation.
  • It doesn’t mean the subgroup-level view is always the “true” one and the aggregate is always “wrong.” Breaking data into subgroups isn’t automatically more truthful, the relevant grouping depends on the question being asked, and controlling for the wrong variable can be just as misleading as aggregating across an important one.
  • It isn’t limited to two subgroups or to medical statistics. It can appear with any number of underlying groups and shows up constantly in business metrics wherever the case mix, the composition of what’s being measured, shifts over time.

Diagnostic Question

Has the mix of what’s being measured changed, and if we hold the mix constant, does the trend still point the same direction?

Explore Further

Field Notes

None yet.

Related Field Guide

Origin

Edward H. Simpson, “The Interpretation of Interaction in Contingency Tables” (1951), Journal of the Royal Statistical Society; related effects were noted earlier by Karl Pearson and Udny Yule, and the pattern is sometimes called the Yule-Simpson effect.

Know someone who’d enjoy this?