[edit]
Estimating Upcoding in Medicare Advantage: Identifying Contributors, Costs, and Mechanisms
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:2227-2256, 2026.
Abstract
Strategic misreporting – the manipulation of features to obtain a better outcome – undermines the integrity of machine learning models used for healthcare resource allocation. In the U.S. Medicare Advantage program, private insurers are reimbursed based on the beneficiary diagnoses they report, creating financial incentives that encourage over-reporting, leading to billions of dollars in annual overpayments. Despite the magnitude of the problem, estimating insurers’ misreporting rates is problematic due to the lack of ground-truth diagnoses. Past work has used causality to try to estimate misreporting rates, but assumes access to an unmanipulated dataset and fails to account for unobserved confounders. To mitigate these issues, we introduce the Stitched Causal Misreporting Estimator (SCaMEr), a causally motivated auditing approach to estimate insurer-specific misreporting rates. SCaMEr addresses two key limitations of prior work: (1) it avoids the need for an unmanipulated reference dataset, and (2) it gives reliable upper and lower bounds on misreporting estimates in the presence of hidden confounders. To overcome the lack of an unmanipulated reference, SCaMEr "stitches" multiple datasets together with varying misreporting rates to recover unbiased estimates. To account for unobserved confounding, it incorporates causal sensitivity analysis to produce uncertainty bounds. We validate SCaMEr on both semi-synthetic and real-world Medicare Advantage data, where it achieves lower estimation error than baselines. Our results show that SCaMEr can enable auditing by identifying diagnoses, insurance plans, and reporting mechanisms that are susceptible to misreporting.