[edit]
The Severity Trap: Benchmarking Sequential Representation Learning for Oxygenation Trajectory Phenotyping in Acute Hypoxaemic Respiratory Failure
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1264-1299, 2026.
Abstract
Patients with acute hypoxaemic respiratory failure (AHRF) presenting with similar early physiology can follow markedly different trajectories, motivating trajectory phenotyping for mechanistic discovery and trial enrichment. Statistical models such as competing-risk latent class mixed models identify clinically meaningful classes but scale poorly to high-dimensional, irregular ICU time series, motivating sequential representation learning with variational autoencoders (VAEs). We characterise a systematic failure mode of learned trajectory representations, the \emph{severity trap}: representations organise primarily along baseline illness severity rather than trajectory shape, even with auxiliary survival objectives or disentanglement constraints. This is an instance of shortcut learning, or nuisance-variable dominance, that we characterise for clinical trajectory phenotyping and show evades the unsupervised metrics commonly used to select such models. Using engineered baselines of simple trajectory features that match, and once given the same outcome signal exceed, every deep model, together with controlled synthetic experiments, we trace the failure to reconstruction-based objectives preferentially encoding the highest-variance factor of the input, here baseline severity. We introduce a multi-cohort benchmark evaluating sequential representation learners (VAE families as the primary object of study; masked-transformer, diffusion, and contrastive encoders as stress tests) against four externally validated oxygenation trajectory archetypes across three international ICU cohorts (MIMIC-IV, Imperial College Healthcare NHS Trust, and Amsterdam UMCdb). Every reconstruction- or level-preserving objective falls into the trap. Only an explicitly level-invariant contrastive objective partially escapes, and clustering stability and reconstruction error correlate little with clinical validity. These findings highlight the need for evaluation grounded in externally validated clinical phenotypes, with nuisance-aware baselines, when developing representation learning for healthcare time series.