Back to Module 3: Quality Measurement

Lesson 3

From Measurement to Improvement

About 4 min

Scores do not improve care by themselves. This lesson covers what separates measurement that moves care from measurement that moves numbers.

Measurement is a means. The practical question that closes this module: under what conditions does measurement improve care, and when does it merely improve scores?

Benchmarks are policy choices

What counts as “good” is itself a design decision:

ApproachStrengthWeakness
Percentile (vs. peers)Always identifies leaders and laggardsCan reward mediocrity in a weak field
Absolute thresholdClear and stableThe bar is arbitrary
Improvement scoringKeeps low performers engagedPays again for ground others already hold

Mature programs blend these, typically scoring the better of attainment or improvement. And wherever benchmarks derive from historical performance, the ratchet returns: this year’s success becomes next year’s harder target.

The gaming spectrum

Any measure attached to money will be optimized. The design question is what kind of optimization you get:

  1. True improvement (outreach staff closing real gaps)
  2. Documentation improvement (care delivered but previously unrecorded)
  3. Effort reallocation (measured tasks crowd out unmeasured care)
  4. Exclusion engineering (working denominator rules so hard patients stop counting)
  5. Selection and misreporting (avoiding patients who hurt scores; falsifying data)

Audited specifications limit misreporting; risk adjustment blunts selection; rotating measure sets limits teaching to the test. No design eliminates optimization. The goal is making the profitable response and the clinically right response the same thing.

What actually moves care

Decades of quality improvement work point to four conditions:

  • Timely, specific feedback. A monthly list of named, overdue patients is a work queue. A stale panel-level score is trivia.
  • Owned work. Someone must be responsible for closing gaps, with registries, standing orders, and protected time.
  • Believed measures. Clinicians who think a measure is clinically wrong comply minimally and game freely. Involve them in selection; drop bad measures.
  • Stakes that fit the signal. Small unreliable differences with large payments attached produce anxiety and gaming, as the Stars and MIPS stories show.

Worth remembering: measurement is one link in an improvement loop: measures find gaps, data systems surface them to people who can act, incentives fund the capacity to act, and evaluation checks that the loop produces health rather than documentation.

Key takeaways

  • Benchmark design determines who wins and whether the middle improves.
  • Gaming is a spectrum; design determines where behavior lands.
  • Feedback, ownership, credibility, and proportionate stakes turn scores into better care.

Check your understanding

A clinician learns in March about her October screening performance, reported at panel level with no patient list. What is the most likely result?

Share