A comparison group only works if it is truly comparable. This lesson covers matching and difference-in-differences in plain language.
A comparison group only helps if it is genuinely comparable. A badly chosen one can be worse than none, because it lends false credibility to a biased result. This lesson covers the two workhorse techniques evaluators use to build and use a fair comparison: matching and difference-in-differences.
Matching: making the groups alike
If program participants are healthier or better-resourced than the general population (the confounding problem), you cannot compare them to everyone else. Matching builds a comparison group that resembles the program group on observable traits: age, diagnoses, prior spending, geography, practice size. The goal is a comparison group that looks like the program group in every measurable way except the program itself.
Matching has a hard limit worth stating plainly:
Worth remembering: matching can only balance what you can measure. If joiners differ on something unrecorded, motivation, management quality, unmeasured health, matching cannot fix it. This is why even well-matched observational studies carry a caveat that randomization would not.
Difference-in-differences: subtracting the background
Once you have a comparison group, the cleanest way to use it is difference-in-differences. The name is literal: take the program group’s change over time, take the comparison group’s change over the same time, and subtract.
| Group | Before | After | Change |
|---|---|---|---|
| Program | $9,000 | $9,300 | +$300 |
| Comparison | $9,000 | $9,700 | +$700 |
| Difference-in-differences | -$400 |
Both groups’ costs rose, because the background trend pushed both up. But the program group rose $400 less. That $400, the difference between the two differences, is the estimated program effect. It automatically nets out any trend that hit both groups equally.
What can still go wrong
Difference-in-differences rests on one big assumption: that without the program, the two groups would have moved in parallel. If the program group was already on a different trajectory before the program (say, its costs were already slowing), the method attributes that pre-existing divergence to the program. Good studies test this by checking that the groups moved together in the years before the program began. When you read an evaluation, look for that check.
Key takeaways
- Matching builds a comparison group that resembles the program group on measurable traits, but cannot balance the unmeasured.
- Difference-in-differences subtracts the comparison group’s change from the program group’s, netting out shared trends.
- It assumes the groups would have moved in parallel absent the program; credible studies show the pre-period trends matched.
Check your understanding
What does a difference-in-differences approach measure?
Difference-in-differences subtracts the comparison group's change (the background trend) from the program group's change, leaving the program's added effect.