[edit]
Auditing Measurement Processes in Clinical Registries: Detecting Label–Environment Confounding Before Model Development
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1834-1864, 2026.
Abstract
Clinical prediction models trained on registry data routinely achieve high in-distribution accuracy, yet their transportability across institutions, clinicians, and time periods remains poorly understood. We formalize this problem through a measurement process model in which the observed label $Y^{\text{obs}}$ depends on both a true clinical state $Y^$*$$ and a reporting indicator $R(e,t)$ that varies with environment $e$ and era $t$: when reporting is inactive, the true label is unobserved and may be imputed by registry convention, introducing structured bias. We propose a measurement audita battery of three permutation-calibrated statistical tests that detects label–environment entanglement before any predictive model is builtand an evaluation-based attribution framework that estimates how much of a model’s apparent discrimination is due to environment confounding versus genuine clinical signal. Applied to a cochlear implant registry of 3,584 patients, the audit identifies severe, degenerate confounding (Cramér’s $V = 0.39$ for label–era association, 4 of 8 eras with zero negative labels, permutation $p<0.002$). The attribution reveals that approximately 10% of the IID AUROC ($0.978$, 95% CI $[0.955,0.991]$) traces to era confounding; the era-blocked AUROC of $0.875$ ($[0.668,0.925]$) better estimates transportable clinical discrimination. External validation on the MIMIC-IV ICU mortality cohort (n=67,224; 7 care-unit environments; 5 eras) shows that the audit distinguishes qualitatively different confounding regimes: it detects diffuse, unit-driven confounding ($V=0.145$, permutation $p<0.002$; confounding share 3.4% of IID AUROC) while correctly passing era-level tests on a registry with no zero-negative strata. A semi-synthetic calibration study confirms favorable operating properties (power rising monotonically with confounding strength; 4% false-positive rate; zero empirical family-wise error under label permutation). We discuss concrete remediation strategies for registries that fail the audit, and propose environment-aware evaluation as a candidate pre-modeling diagnostic for registry-trained clinical ML.