A result that is true in one place, time, and population may not travel. This lesson covers when evidence transfers and when it does not.
An evaluation answers a question about a specific program, in a specific population, at a specific time. Whether its answer travels to a different program, population, or era is a separate question, and assuming it does is a quiet, common error. This is generalizability, sometimes called external validity.
Why results do not always travel
A finding is a product of its context, and several parts of that context may not repeat:
- Population. A model that worked for a stable, commercially insured population may not work for a churning Medicaid one (the Medicaid course made this concrete).
- Condition. Bundled payment succeeded for elective joint replacement, a predictable, standardizable episode. That is weak evidence it will work for heart failure, which starts unpredictably and varies widely, exactly what the case-studies module found.
- Era. A result from an era of high medical trend may not hold when trend is low, because the savings opportunity itself changed.
- Scale and selection. A voluntary pilot of eager early adopters may not predict what happens when a model is mandatory and universal.
Worth remembering: the right question is never “did it work?” but “did it work, for whom, and would those conditions hold here?” A result proven in one context is a hypothesis in another, not a conclusion.
The early-adopter trap
Generalizability interacts with selection. The organizations that pilot a new model are usually the most capable and motivated. Even a rigorously evaluated success among early adopters may not survive contact with the average organization, which has less capital, weaker data, and less appetite for risk. Scaling a model tends to dilute the results the pilot reported, one reason models that shine in demonstration disappoint at national scale.
How to read for it
- Note the population, condition, and era the result came from, and how far they are from the setting you care about.
- Distinguish efficacy (it worked under favorable pilot conditions) from effectiveness (it works in the messy real world at scale).
- Treat a single positive result as a reason to test further in the new context, not as settled proof it will transfer.
Key takeaways
- Results are products of a specific population, condition, and era, and may not transfer to others.
- Evidence generalizes best to similar settings; surgical-bundle success does not imply medical-bundle success.
- Early-adopter pilots tend to overstate what a model achieves at scale, so treat transfer as a hypothesis to test.
Check your understanding
A bundled-payment model saved money for elective joint replacements. What is the safest conclusion?
Evidence generalizes best to similar conditions, populations, and eras. Elective surgical bundles worked partly because they are predictable and standardizable, which does not describe medical or emergency care.