Two competing conclusions about the exact same dataset can both be mathematically true. This isn't a trick. It's Simpson's paradox, and it explains why your data might be sabotaging your judgment without either of you noticing.
Most people assume that if something is true in all the parts, it must be true in the whole. If a medication works better than a placebo in men and works better in women, then obviously it works better overall, right? Not necessarily. According to research on statistical phenomena documented at the Stanford Encyclopedia of Philosophy, Simpson's paradox is a phenomenon where a trend appears in multiple separate data groups but reverses—or vanishes entirely—when those groups are combined. The math isn't wrong either way. Both conclusions are genuine. One dataset, two contradictory truths.
The classic example involves a medical treatment studied across two hospitals. At Hospital A, the treatment has a higher success rate than the control in both male and female patients. At Hospital B, the same pattern holds: the treatment outperforms the control for both men and women. But when you combine all patients from both hospitals, the control suddenly looks better. How? The answer lurks in what statisticians call "confounding variables"—hidden factors that change the composition of your groups. In this case, Hospital A might treat sicker patients overall, or Hospital B might have enrolled more women, and the treatment and control might work differently across gender lines. The individual trends are real. The reversal is real. The paradox is that your eyes are telling the truth in two incompatible ways.
The phenomenon shows up everywhere real institutions actually make decisions. A 2020 analysis published in the National Center for Biotechnology Information examined how Simpson's paradox has appeared in medical treatment evaluations, hiring bias investigations, and educational research. One infamous case involved UC Berkeley's graduate admissions data in the 1970s. Overall, the university appeared to discriminate against women applicants—men had a higher admission rate. But when analyzed by individual department, women actually had equal or slightly better acceptance rates in most programs. The reversal happened because women applied disproportionately to competitive departments with lower acceptance rates overall, while men clustered in departments that accepted more applicants. Both statistics were real. Neither was a lie. Yet they pointed in opposite directions.
Why does this happen? Because aggregation hides the structure of your data. When you combine groups with different sizes, different compositions, or different underlying distributions, you're essentially adding weights to your conclusions that aren't visible in the final number. The treatment that looks better in every subgroup might be applied primarily to a population where outcomes are naturally worse anyway—the effect exists, but the denominator effect (the base rate in each group) overwhelms the actual treatment benefit when you zoom out. The paradox isn't a statistical error. It's a reminder that "which groups do we look at?" is a more consequential question than "what does the data show?"
The real danger isn't the paradox itself—it's that it's entirely possible to be intellectually honest and still draw opposite conclusions from the same data by choosing which level of aggregation to examine. A company might truthfully show that salary increases went to more women this year (subgroup truth) while still underpaying women overall (aggregate truth). A medication could genuinely work better in all demographics yet contribute to worse health outcomes in a population if it's being prescribed to the sickest patients. The paradox is a gift to anyone motivated to prove their point: the math won't stop you.