Claims are the backbone of value-based analytics. This lesson covers what they are, how they are structured, and the quirks that trip up analysts.
Claims data is the backbone of value-based analytics: the source that supports cost, utilization, attribution, and risk scoring. Understanding what it is, how it is structured, and where it misleads is the foundation of building analytics for value-based care.
What claims data is
A claim is a billing record: the request a provider sends a payer for payment, and the payer’s adjudication of it. Claims are standardized around code sets the earlier courses referenced, ICD-10 diagnoses, CPT and HCPCS procedures, NDC drug codes, which makes them consistent and analyzable across providers and payers.
Their defining virtue is completeness across settings. Because every paid encounter generates a claim, the data captures care wherever the patient received it: the hospital across town, the specialist you never referred to, the filled prescription. That whole-population, all-settings view is what makes claims the primary source for population analytics.
How claims are structured
Analysts work with a recognizable structure:
- Header and line items. A claim has header-level information (patient, provider, dates, total) and line items detailing each service, code, and amount.
- Institutional versus professional. Facility claims (hospitals, institutions) and professional claims (physicians) have different formats and fields.
- Adjudication detail. Billed, allowed, and paid amounts differ; denied lines, adjustments, and the paid amount all matter for accurate cost.
The quirks that trip up analysts
Worth remembering: claims data is complete but lagging and clinically thin, and both quirks bite analysts specifically. The claims lag from the finance course means recent data is incomplete, so any current-period analysis needs completion factors, not raw paid claims. And claims record what was billed, not what happened clinically: a claim shows an A1c test was performed, never the result. Analysts who treat claims as clinical truth, or recent claims as complete, produce confident, wrong numbers.
A third caveat, from the risk-adjustment course: claims reflect coding behavior as much as clinical reality, so diagnosis-based measures inherit coding-intensity effects.
Working with claims
Practically, building on claims means ingesting standardized files (often in defined formats), handling the adjudication and runout, and joining claims into episodes and patient histories. Done well, claims become the substrate for nearly every value-based metric. Done carelessly, their lag and billing-not-clinical nature quietly corrupt the analysis. The rest of this course is partly about combining claims with the clinical and social data that fill their gaps.
Key takeaways
- Claims are standardized billing records, complete across all settings, and the backbone of population analytics.
- They are structured as headers and line items, with billed, allowed, and paid amounts that must be handled correctly.
- They lag, are clinically thin, and reflect billing behavior, so analysts must use completion factors and never mistake claims for clinical truth.
Check your understanding
What makes claims data uniquely valuable for population analytics?
Because claims follow the patient across all settings and providers, they show the whole picture of utilization and spending, the visibility no single organization's records provide.