The Severity Trap: Benchmarking Sequential Representation Learning for Oxygenation Trajectory Phenotyping in Acute Hypoxaemic Respiratory Failure

Dominic C Marshall, Nikita Narayanan, David B Antcliffe, Sonali Parbhoo
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1264-1299, 2026.

Abstract

Patients with acute hypoxaemic respiratory failure (AHRF) presenting with similar early physiology can follow markedly different trajectories, motivating trajectory phenotyping for mechanistic discovery and trial enrichment. Statistical models such as competing-risk latent class mixed models identify clinically meaningful classes but scale poorly to high-dimensional, irregular ICU time series, motivating sequential representation learning with variational autoencoders (VAEs). We characterise a systematic failure mode of learned trajectory representations, the \emph{severity trap}: representations organise primarily along baseline illness severity rather than trajectory shape, even with auxiliary survival objectives or disentanglement constraints. This is an instance of shortcut learning, or nuisance-variable dominance, that we characterise for clinical trajectory phenotyping and show evades the unsupervised metrics commonly used to select such models. Using engineered baselines of simple trajectory features that match, and once given the same outcome signal exceed, every deep model, together with controlled synthetic experiments, we trace the failure to reconstruction-based objectives preferentially encoding the highest-variance factor of the input, here baseline severity. We introduce a multi-cohort benchmark evaluating sequential representation learners (VAE families as the primary object of study; masked-transformer, diffusion, and contrastive encoders as stress tests) against four externally validated oxygenation trajectory archetypes across three international ICU cohorts (MIMIC-IV, Imperial College Healthcare NHS Trust, and Amsterdam UMCdb). Every reconstruction- or level-preserving objective falls into the trap. Only an explicitly level-invariant contrastive objective partially escapes, and clustering stability and reconstruction error correlate little with clinical validity. These findings highlight the need for evaluation grounded in externally validated clinical phenotypes, with nuisance-aware baselines, when developing representation learning for healthcare time series.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-marshall26a, title = {The Severity Trap: Benchmarking Sequential Representation Learning for Oxygenation Trajectory Phenotyping in Acute Hypoxaemic Respiratory Failure}, author = {Marshall, Dominic C and Narayanan, Nikita and Antcliffe, David B and Parbhoo, Sonali}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1264--1299}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/marshall26a/marshall26a.pdf}, url = {https://proceedings.mlr.press/v340/marshall26a.html}, abstract = {Patients with acute hypoxaemic respiratory failure (AHRF) presenting with similar early physiology can follow markedly different trajectories, motivating trajectory phenotyping for mechanistic discovery and trial enrichment. Statistical models such as competing-risk latent class mixed models identify clinically meaningful classes but scale poorly to high-dimensional, irregular ICU time series, motivating sequential representation learning with variational autoencoders (VAEs). We characterise a systematic failure mode of learned trajectory representations, the \emph{severity trap}: representations organise primarily along baseline illness severity rather than trajectory shape, even with auxiliary survival objectives or disentanglement constraints. This is an instance of shortcut learning, or nuisance-variable dominance, that we characterise for clinical trajectory phenotyping and show evades the unsupervised metrics commonly used to select such models. Using engineered baselines of simple trajectory features that match, and once given the same outcome signal exceed, every deep model, together with controlled synthetic experiments, we trace the failure to reconstruction-based objectives preferentially encoding the highest-variance factor of the input, here baseline severity. We introduce a multi-cohort benchmark evaluating sequential representation learners (VAE families as the primary object of study; masked-transformer, diffusion, and contrastive encoders as stress tests) against four externally validated oxygenation trajectory archetypes across three international ICU cohorts (MIMIC-IV, Imperial College Healthcare NHS Trust, and Amsterdam UMCdb). Every reconstruction- or level-preserving objective falls into the trap. Only an explicitly level-invariant contrastive objective partially escapes, and clustering stability and reconstruction error correlate little with clinical validity. These findings highlight the need for evaluation grounded in externally validated clinical phenotypes, with nuisance-aware baselines, when developing representation learning for healthcare time series.} }
Endnote
%0 Conference Paper %T The Severity Trap: Benchmarking Sequential Representation Learning for Oxygenation Trajectory Phenotyping in Acute Hypoxaemic Respiratory Failure %A Dominic C Marshall %A Nikita Narayanan %A David B Antcliffe %A Sonali Parbhoo %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-marshall26a %I PMLR %P 1264--1299 %U https://proceedings.mlr.press/v340/marshall26a.html %V 340 %X Patients with acute hypoxaemic respiratory failure (AHRF) presenting with similar early physiology can follow markedly different trajectories, motivating trajectory phenotyping for mechanistic discovery and trial enrichment. Statistical models such as competing-risk latent class mixed models identify clinically meaningful classes but scale poorly to high-dimensional, irregular ICU time series, motivating sequential representation learning with variational autoencoders (VAEs). We characterise a systematic failure mode of learned trajectory representations, the \emph{severity trap}: representations organise primarily along baseline illness severity rather than trajectory shape, even with auxiliary survival objectives or disentanglement constraints. This is an instance of shortcut learning, or nuisance-variable dominance, that we characterise for clinical trajectory phenotyping and show evades the unsupervised metrics commonly used to select such models. Using engineered baselines of simple trajectory features that match, and once given the same outcome signal exceed, every deep model, together with controlled synthetic experiments, we trace the failure to reconstruction-based objectives preferentially encoding the highest-variance factor of the input, here baseline severity. We introduce a multi-cohort benchmark evaluating sequential representation learners (VAE families as the primary object of study; masked-transformer, diffusion, and contrastive encoders as stress tests) against four externally validated oxygenation trajectory archetypes across three international ICU cohorts (MIMIC-IV, Imperial College Healthcare NHS Trust, and Amsterdam UMCdb). Every reconstruction- or level-preserving objective falls into the trap. Only an explicitly level-invariant contrastive objective partially escapes, and clustering stability and reconstruction error correlate little with clinical validity. These findings highlight the need for evaluation grounded in externally validated clinical phenotypes, with nuisance-aware baselines, when developing representation learning for healthcare time series.
APA
Marshall, D.C., Narayanan, N., Antcliffe, D.B. & Parbhoo, S.. (2026). The Severity Trap: Benchmarking Sequential Representation Learning for Oxygenation Trajectory Phenotyping in Acute Hypoxaemic Respiratory Failure. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1264-1299 Available from https://proceedings.mlr.press/v340/marshall26a.html.

Related Material