Estimating Upcoding in Medicare Advantage: Identifying Contributors, Costs, and Mechanisms

Dylan Zapzalka, Muskaan Mittal, Jenna Wiens, Maggie Makar
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:2227-2256, 2026.

Abstract

Strategic misreporting – the manipulation of features to obtain a better outcome – undermines the integrity of machine learning models used for healthcare resource allocation. In the U.S. Medicare Advantage program, private insurers are reimbursed based on the beneficiary diagnoses they report, creating financial incentives that encourage over-reporting, leading to billions of dollars in annual overpayments. Despite the magnitude of the problem, estimating insurers’ misreporting rates is problematic due to the lack of ground-truth diagnoses. Past work has used causality to try to estimate misreporting rates, but assumes access to an unmanipulated dataset and fails to account for unobserved confounders. To mitigate these issues, we introduce the Stitched Causal Misreporting Estimator (SCaMEr), a causally motivated auditing approach to estimate insurer-specific misreporting rates. SCaMEr addresses two key limitations of prior work: (1) it avoids the need for an unmanipulated reference dataset, and (2) it gives reliable upper and lower bounds on misreporting estimates in the presence of hidden confounders. To overcome the lack of an unmanipulated reference, SCaMEr "stitches" multiple datasets together with varying misreporting rates to recover unbiased estimates. To account for unobserved confounding, it incorporates causal sensitivity analysis to produce uncertainty bounds. We validate SCaMEr on both semi-synthetic and real-world Medicare Advantage data, where it achieves lower estimation error than baselines. Our results show that SCaMEr can enable auditing by identifying diagnoses, insurance plans, and reporting mechanisms that are susceptible to misreporting.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-zapzalka26a, title = {Estimating Upcoding in Medicare Advantage: Identifying Contributors, Costs, and Mechanisms}, author = {Zapzalka, Dylan and Mittal, Muskaan and Wiens, Jenna and Makar, Maggie}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {2227--2256}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/zapzalka26a/zapzalka26a.pdf}, url = {https://proceedings.mlr.press/v340/zapzalka26a.html}, abstract = {Strategic misreporting – the manipulation of features to obtain a better outcome – undermines the integrity of machine learning models used for healthcare resource allocation. In the U.S. Medicare Advantage program, private insurers are reimbursed based on the beneficiary diagnoses they report, creating financial incentives that encourage over-reporting, leading to billions of dollars in annual overpayments. Despite the magnitude of the problem, estimating insurers’ misreporting rates is problematic due to the lack of ground-truth diagnoses. Past work has used causality to try to estimate misreporting rates, but assumes access to an unmanipulated dataset and fails to account for unobserved confounders. To mitigate these issues, we introduce the Stitched Causal Misreporting Estimator (SCaMEr), a causally motivated auditing approach to estimate insurer-specific misreporting rates. SCaMEr addresses two key limitations of prior work: (1) it avoids the need for an unmanipulated reference dataset, and (2) it gives reliable upper and lower bounds on misreporting estimates in the presence of hidden confounders. To overcome the lack of an unmanipulated reference, SCaMEr "stitches" multiple datasets together with varying misreporting rates to recover unbiased estimates. To account for unobserved confounding, it incorporates causal sensitivity analysis to produce uncertainty bounds. We validate SCaMEr on both semi-synthetic and real-world Medicare Advantage data, where it achieves lower estimation error than baselines. Our results show that SCaMEr can enable auditing by identifying diagnoses, insurance plans, and reporting mechanisms that are susceptible to misreporting.} }
Endnote
%0 Conference Paper %T Estimating Upcoding in Medicare Advantage: Identifying Contributors, Costs, and Mechanisms %A Dylan Zapzalka %A Muskaan Mittal %A Jenna Wiens %A Maggie Makar %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-zapzalka26a %I PMLR %P 2227--2256 %U https://proceedings.mlr.press/v340/zapzalka26a.html %V 340 %X Strategic misreporting – the manipulation of features to obtain a better outcome – undermines the integrity of machine learning models used for healthcare resource allocation. In the U.S. Medicare Advantage program, private insurers are reimbursed based on the beneficiary diagnoses they report, creating financial incentives that encourage over-reporting, leading to billions of dollars in annual overpayments. Despite the magnitude of the problem, estimating insurers’ misreporting rates is problematic due to the lack of ground-truth diagnoses. Past work has used causality to try to estimate misreporting rates, but assumes access to an unmanipulated dataset and fails to account for unobserved confounders. To mitigate these issues, we introduce the Stitched Causal Misreporting Estimator (SCaMEr), a causally motivated auditing approach to estimate insurer-specific misreporting rates. SCaMEr addresses two key limitations of prior work: (1) it avoids the need for an unmanipulated reference dataset, and (2) it gives reliable upper and lower bounds on misreporting estimates in the presence of hidden confounders. To overcome the lack of an unmanipulated reference, SCaMEr "stitches" multiple datasets together with varying misreporting rates to recover unbiased estimates. To account for unobserved confounding, it incorporates causal sensitivity analysis to produce uncertainty bounds. We validate SCaMEr on both semi-synthetic and real-world Medicare Advantage data, where it achieves lower estimation error than baselines. Our results show that SCaMEr can enable auditing by identifying diagnoses, insurance plans, and reporting mechanisms that are susceptible to misreporting.
APA
Zapzalka, D., Mittal, M., Wiens, J. & Makar, M.. (2026). Estimating Upcoding in Medicare Advantage: Identifying Contributors, Costs, and Mechanisms. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:2227-2256 Available from https://proceedings.mlr.press/v340/zapzalka26a.html.

Related Material