The single most common way health programs fool themselves. This lesson explains why targeting last year's most expensive patients guarantees a false success.
If you learn only one trap from this course, make it this one. Regression to the mean has produced more false success stories in health care than any other statistical mistake, and it is built into the way value-based programs pick their patients.
The phenomenon
Any measure that bounces around from year to year has a simple property: a group selected for being extreme at one point tends to be less extreme the next, on its own, with no intervention at all. The tallest kids in a class are, on average, closer to average height when re-measured against a different group; a baseball player’s best month is usually followed by a more ordinary one. Not because anything acted on them, but because the extreme was partly luck, and luck does not repeat.
Medical cost is exactly this kind of measure. A patient’s costliest year is often driven by a one-time event, a surgery, a crisis, a bad break, that will not recur. Next year, on average, they cost less. Automatically.
Why it wrecks program evaluation
Value-based programs love to target the highest-cost patients, and for good clinical reasons. But that targeting is the trap:
Worth remembering: select the top few percent of spenders this year, do absolutely nothing, and their average cost will fall next year through regression to the mean alone. A program that enrolls last year’s most expensive patients and reports that their costs dropped has not shown that it worked. It has shown that arithmetic works.
The intro course’s population-health module flagged this: last year’s high utilizers partially regress on their own. Here is the evaluation consequence spelled out. The “savings” a care-management program claims for its high-risk enrollees are, in part or in whole, regression to the mean wearing a lab coat.
How to catch it
There is one reliable defense, and it is the theme of this course: a comparison group selected the same way. Take last year’s highest-cost patients, and only enroll half of them (or compare to a matched high-cost group elsewhere). Both groups will regress. If the enrolled group falls more than the comparison group, that extra drop is a candidate for the real effect. Without that comparison, a favorable result for a high-cost cohort tells you nothing.
When you read that a program “reduced costs for its highest-risk members,” your immediate question should be: compared to a similar high-risk group that was not enrolled? If the study cannot answer that, regression to the mean is the most likely explanation for the result.
Key takeaways
- Groups selected for being extreme drift back toward average on their own, with no intervention.
- Targeting last year’s highest-cost patients builds this drift directly into a program’s results.
- Only a comparison group selected the same way separates a real effect from regression to the mean.
Check your understanding
A program enrolls the highest-cost patients from last year, and their costs fall this year. Why is that not proof the program worked?
Regression to the mean: a group selected for being extreme this year will drift toward average next year regardless of any intervention. Without a comparison group, that natural drift looks like a program effect.