Auditing Measurement Processes in Clinical Registries: Detecting Label–Environment Confounding Before Model Development

Apurv Shukla, Pravansu Mohanty, Devin L McCaslin
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1834-1864, 2026.

Abstract

Clinical prediction models trained on registry data routinely achieve high in-distribution accuracy, yet their transportability across institutions, clinicians, and time periods remains poorly understood. We formalize this problem through a measurement process model in which the observed label $Y^{\text{obs}}$ depends on both a true clinical state $Y^$*$$ and a reporting indicator $R(e,t)$ that varies with environment $e$ and era $t$: when reporting is inactive, the true label is unobserved and may be imputed by registry convention, introducing structured bias. We propose a measurement audita battery of three permutation-calibrated statistical tests that detects label–environment entanglement before any predictive model is builtand an evaluation-based attribution framework that estimates how much of a model’s apparent discrimination is due to environment confounding versus genuine clinical signal. Applied to a cochlear implant registry of 3,584 patients, the audit identifies severe, degenerate confounding (Cramér’s $V = 0.39$ for label–era association, 4 of 8 eras with zero negative labels, permutation $p<0.002$). The attribution reveals that approximately 10% of the IID AUROC ($0.978$, 95% CI $[0.955,0.991]$) traces to era confounding; the era-blocked AUROC of $0.875$ ($[0.668,0.925]$) better estimates transportable clinical discrimination. External validation on the MIMIC-IV ICU mortality cohort (n=67,224; 7 care-unit environments; 5 eras) shows that the audit distinguishes qualitatively different confounding regimes: it detects diffuse, unit-driven confounding ($V=0.145$, permutation $p<0.002$; confounding share 3.4% of IID AUROC) while correctly passing era-level tests on a registry with no zero-negative strata. A semi-synthetic calibration study confirms favorable operating properties (power rising monotonically with confounding strength; 4% false-positive rate; zero empirical family-wise error under label permutation). We discuss concrete remediation strategies for registries that fail the audit, and propose environment-aware evaluation as a candidate pre-modeling diagnostic for registry-trained clinical ML.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-shukla26a, title = {Auditing Measurement Processes in Clinical Registries: Detecting Label–Environment Confounding Before Model Development}, author = {Shukla, Apurv and Mohanty, Pravansu and McCaslin, Devin L}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1834--1864}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/shukla26a/shukla26a.pdf}, url = {https://proceedings.mlr.press/v340/shukla26a.html}, abstract = {Clinical prediction models trained on registry data routinely achieve high in-distribution accuracy, yet their transportability across institutions, clinicians, and time periods remains poorly understood. We formalize this problem through a measurement process model in which the observed label $Y^{\text{obs}}$ depends on both a true clinical state $Y^$*$$ and a reporting indicator $R(e,t)$ that varies with environment $e$ and era $t$: when reporting is inactive, the true label is unobserved and may be imputed by registry convention, introducing structured bias. We propose a measurement audita battery of three permutation-calibrated statistical tests that detects label–environment entanglement before any predictive model is builtand an evaluation-based attribution framework that estimates how much of a model’s apparent discrimination is due to environment confounding versus genuine clinical signal. Applied to a cochlear implant registry of 3,584 patients, the audit identifies severe, degenerate confounding (Cramér’s $V = 0.39$ for label–era association, 4 of 8 eras with zero negative labels, permutation $p<0.002$). The attribution reveals that approximately 10% of the IID AUROC ($0.978$, 95% CI $[0.955,0.991]$) traces to era confounding; the era-blocked AUROC of $0.875$ ($[0.668,0.925]$) better estimates transportable clinical discrimination. External validation on the MIMIC-IV ICU mortality cohort (n=67,224; 7 care-unit environments; 5 eras) shows that the audit distinguishes qualitatively different confounding regimes: it detects diffuse, unit-driven confounding ($V=0.145$, permutation $p<0.002$; confounding share 3.4% of IID AUROC) while correctly passing era-level tests on a registry with no zero-negative strata. A semi-synthetic calibration study confirms favorable operating properties (power rising monotonically with confounding strength; 4% false-positive rate; zero empirical family-wise error under label permutation). We discuss concrete remediation strategies for registries that fail the audit, and propose environment-aware evaluation as a candidate pre-modeling diagnostic for registry-trained clinical ML.} }
Endnote
%0 Conference Paper %T Auditing Measurement Processes in Clinical Registries: Detecting Label–Environment Confounding Before Model Development %A Apurv Shukla %A Pravansu Mohanty %A Devin L McCaslin %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-shukla26a %I PMLR %P 1834--1864 %U https://proceedings.mlr.press/v340/shukla26a.html %V 340 %X Clinical prediction models trained on registry data routinely achieve high in-distribution accuracy, yet their transportability across institutions, clinicians, and time periods remains poorly understood. We formalize this problem through a measurement process model in which the observed label $Y^{\text{obs}}$ depends on both a true clinical state $Y^$*$$ and a reporting indicator $R(e,t)$ that varies with environment $e$ and era $t$: when reporting is inactive, the true label is unobserved and may be imputed by registry convention, introducing structured bias. We propose a measurement audita battery of three permutation-calibrated statistical tests that detects label–environment entanglement before any predictive model is builtand an evaluation-based attribution framework that estimates how much of a model’s apparent discrimination is due to environment confounding versus genuine clinical signal. Applied to a cochlear implant registry of 3,584 patients, the audit identifies severe, degenerate confounding (Cramér’s $V = 0.39$ for label–era association, 4 of 8 eras with zero negative labels, permutation $p<0.002$). The attribution reveals that approximately 10% of the IID AUROC ($0.978$, 95% CI $[0.955,0.991]$) traces to era confounding; the era-blocked AUROC of $0.875$ ($[0.668,0.925]$) better estimates transportable clinical discrimination. External validation on the MIMIC-IV ICU mortality cohort (n=67,224; 7 care-unit environments; 5 eras) shows that the audit distinguishes qualitatively different confounding regimes: it detects diffuse, unit-driven confounding ($V=0.145$, permutation $p<0.002$; confounding share 3.4% of IID AUROC) while correctly passing era-level tests on a registry with no zero-negative strata. A semi-synthetic calibration study confirms favorable operating properties (power rising monotonically with confounding strength; 4% false-positive rate; zero empirical family-wise error under label permutation). We discuss concrete remediation strategies for registries that fail the audit, and propose environment-aware evaluation as a candidate pre-modeling diagnostic for registry-trained clinical ML.
APA
Shukla, A., Mohanty, P. & McCaslin, D.L.. (2026). Auditing Measurement Processes in Clinical Registries: Detecting Label–Environment Confounding Before Model Development. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1834-1864 Available from https://proceedings.mlr.press/v340/shukla26a.html.

Related Material