Back to Module 4: Specialists Inside Population Models

Lesson 3

Measuring Specialist Performance

About 5 min

You cannot steer volume toward better specialists without knowing which ones they are. Specialist measurement is harder than primary care measurement for reasons worth understanding.

Every strategy in this module depends on knowing which specialists produce better results at lower cost. That measurement is genuinely difficult, and organizations that skip the difficulty end up steering volume on the basis of noise.

The four problems

Small numbers. A surgeon may perform forty of a given procedure in a year. Complication rates built on forty cases are dominated by chance, and apparent differences between surgeons frequently vanish on remeasurement. The evaluation course covered regression to the mean; specialist scorecards are where it does the most damage, because a clinician penalized for a bad year will usually look average the next year regardless of any intervention.

Attribution. A patient’s outcome after surgery reflects the surgeon, the anesthesiologist, the hospital, the post-acute facility, and the patient’s own health. Assigning it to one clinician requires a convention, and every convention is arguable. This is the same objection MedPAC raised against MIPS, and it is sharper for proceduralists who work in teams.

Case mix. The specialists who take the hardest cases have the worst raw outcomes. Without adequate risk adjustment, a measurement program penalizes exactly the clinicians a system most needs, and teaches them to decline complex referrals. The risk adjustment module of the introductory course applies directly.

Cost attribution boundaries. Measuring a specialist’s cost requires deciding which spending belongs to them. The episode design choices from Module 1 reappear here: what triggers the specialist’s accountability, how long it lasts, and what is included.

What actually works

Given those constraints, the measures that hold up share properties:

  • Aggregate across time and procedures to build usable denominators, accepting less granularity in exchange for less noise.
  • Measure at the group level where individual volume is too low, which is also where referral decisions can realistically be directed.
  • Prefer process and utilization measures with large denominators over rare outcomes. Imaging rates before back surgery, use of ambulatory surgery centers for appropriate cases, and follow-up visit patterns are measurable with the volumes actually available.
  • Report total episode cost rather than the specialist’s own billing, since the specialist’s professional fee is a small fraction of what their decisions commit.

That last point is the one most often missed. A cardiologist’s own charges tell you almost nothing. The catheterizations they order, the facilities they use, and the admissions that follow tell you nearly everything.

Worth remembering: the honest position is that specialist measurement is good enough to distinguish clear outliers and rarely good enough to rank the middle. That has a practical consequence: use it to identify and address the tails, and use referral relationships, pathway participation, and communication quality to differentiate everyone else. Building an elaborate ranking of the middle of the distribution and steering volume on it will mostly move patients around based on statistical noise, while damaging relationships with clinicians who can see that the ranking is not real.

Key takeaways

  • Small denominators, attribution, case mix, and cost boundaries all make specialist measurement harder than primary care measurement.
  • Aggregating across time and measuring at group level buys usable denominators.
  • Process and utilization measures with large denominators are more reliable than rare outcomes.
  • Total episode cost reflects a specialist’s influence far better than their own billing does.
  • Measurement is generally reliable at the tails and unreliable in the middle, which should shape how it is used.

Sources

Check your understanding

Why is small-numbers a more severe problem for specialist measurement than for primary care measurement?

Share