What 238 Million Medicaid Billing Rows Can Support
A four-step audit of the public file, the exclusion label, billing-pattern comparisons, and a prospective screening experiment. Evidence status varies by step.
A practical economics lab for testing what public healthcare data can support, where a result breaks, and what evidence would make it decision-ready.
Health-policy screening systems allocate investigative attention, create burdens for providers, and can affect access to care. The economic question is whether a score improves a real decision after label error, missing clinical context, false positives, and enforcement selection are counted.
The named case is the HHS Medicaid Provider Spending by HCPCS dataset, released on February 9, 2026. It aggregates 2018 to 2024 billing-provider, servicing-provider, procedure, and month records. HHS warns that state variation may reflect policy, coding, or submission differences and that low-volume rows are suppressed.
The four-part series turns that data release into a competency demonstration: inventory the available fields, define the label, inspect billing patterns, then compare a prospective classifier with simple baselines. Exact label shares, provider comparisons, and classifier metrics are currently marked under reconciliation. They remain visible as validation targets, not established findings.
Health policy analysis sits at the intersection of several challenges: large administrative datasets with limited variables, strong political incentives to find dramatic results, and real consequences when analysis goes wrong. Providers flagged as suspicious may lose their ability to serve vulnerable populations. Genuine fraud may go undetected when screening methods optimize for volume rather than behavioral anomalies.
The working rule is simple: document the decision, population, label, observation window, threshold, and cost of error before interpreting model performance. Each article carries its own evidence status so a public-source fact cannot be confused with an article-reported result.
Our health policy research draws on:
What is documented, what is article-reported, and what still needs a rerun
HHS documents a 3.5 GB provider-spending file covering January 2018 through December 2024, with privacy suppression and state-data-quality cautions.
The series reports that 40% of matched LEIE exclusions are not fraud-related. The matched-label denominator has not yet been publicly reproduced.
Billing volume can reflect service intensity, specialty, access, and patient need. It cannot establish fraud without clinical and investigative context.
The article reports prospective model metrics and baseline comparisons. A frozen sample, label version, and rerun are still required.
Medicaid fraud detection uses billing and provider records to prioritize claims or providers for review. Billing volume alone cannot establish fraud. This site reports a 40% non-fraud share among matched LEIE exclusions, but that estimate is article-reported and not yet publicly reproduced.
The HHS provider-spending dataset aggregates billing-provider, servicing-provider, procedure, and month records from 2018 through 2024. It can support utilization and outlier analysis, but it cannot establish fraud without richer clinical and investigative context.
The research treats fraud screening as an economic decision problem: define the label and threshold, measure false-positive and false-negative costs, compare simple baselines, test out of time, and document who bears each error.
Our health policy research organized by focus area
Examining what public Medicaid data can and cannot tell us about program integrity and provider behavior.
Data validation, classification approaches, and methodological considerations for health policy research.
Future health policy research directions including healthcare access, cost-effectiveness analysis, and program evaluation.
A four-step audit of the public file, the exclusion label, billing-pattern comparisons, and a prospective screening experiment. Evidence status varies by step.
Status: article-reported. The LEIE taxonomy mixes fraud-related convictions with other statutory exclusion reasons; the reported 40% share still needs a matched public output.
Status: under reconciliation. A methods guide to specialty context, observation windows, label contamination, and article-reported provider comparisons.
Status: under reconciliation. A prospective-validation design with article-reported metrics that require one frozen sample and label version.
New health policy analysis delivered directly. No spam, unsubscribe anytime.
Subscribe for Free