Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs

Niklas Raehse, Luregn J Schlapbach, Daphné Chopard
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1613-1666, 2026.

Abstract

Antimicrobial stewardship (AMS) is critical for reducing unnecessary antibiotic exposure, particularly in pediatric intensive care units (PICUs), where clinical uncertainty leads to frequent broad-spectrum use, and overuse amplifies antimicrobial resistance which have long-term consequences. Machine learning has been proposed to support AMS by identifying patient-level opportunities for intervention from electronic health record data. However, prior work focuses on adult populations and predominantly relies on static tabular representations, leaving open questions about target design, temporal modeling, and generalizability in pediatric settings. In this work, we present a systematic benchmarking study of AMS intervention prediction in the PICU. Using the publicly available Paediatric Intensive Care database from China and a private PICU cohort from the University Children’s Hospital Zurich, Switzerland, we define four clinically relevant proxy targets for reducing antibiotic use: intravenous-to-oral switching, de-escalation, discontinuation, and short-course therapy. We then compare tabular, sequence-based, and graph-based temporal models under a unified evaluation framework. We find that model performance is primarily driven by target prevalence and data characteristics rather than model complexity. Sequence models provide improvements in precision-recall trade-off over tabular approaches at coarse (24-hour) resolution, with limited additional gains when finer temporal structure is incorporated. However, these gains come at the cost of poorer calibration, with simpler tabular models producing more reliable probability estimates. Our results highlight the importance of target selection, temporal representation, and calibration in clinical machine learning, and provide practical guidance for developing reliable decision support systems for AMS in pediatric critical care.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-raehse26a, title = {Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs}, author = {Raehse, Niklas and Schlapbach, Luregn J and Chopard, Daphn\'{e}}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1613--1666}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/raehse26a/raehse26a.pdf}, url = {https://proceedings.mlr.press/v340/raehse26a.html}, abstract = {Antimicrobial stewardship (AMS) is critical for reducing unnecessary antibiotic exposure, particularly in pediatric intensive care units (PICUs), where clinical uncertainty leads to frequent broad-spectrum use, and overuse amplifies antimicrobial resistance which have long-term consequences. Machine learning has been proposed to support AMS by identifying patient-level opportunities for intervention from electronic health record data. However, prior work focuses on adult populations and predominantly relies on static tabular representations, leaving open questions about target design, temporal modeling, and generalizability in pediatric settings. In this work, we present a systematic benchmarking study of AMS intervention prediction in the PICU. Using the publicly available Paediatric Intensive Care database from China and a private PICU cohort from the University Children’s Hospital Zurich, Switzerland, we define four clinically relevant proxy targets for reducing antibiotic use: intravenous-to-oral switching, de-escalation, discontinuation, and short-course therapy. We then compare tabular, sequence-based, and graph-based temporal models under a unified evaluation framework. We find that model performance is primarily driven by target prevalence and data characteristics rather than model complexity. Sequence models provide improvements in precision-recall trade-off over tabular approaches at coarse (24-hour) resolution, with limited additional gains when finer temporal structure is incorporated. However, these gains come at the cost of poorer calibration, with simpler tabular models producing more reliable probability estimates. Our results highlight the importance of target selection, temporal representation, and calibration in clinical machine learning, and provide practical guidance for developing reliable decision support systems for AMS in pediatric critical care.} }
Endnote
%0 Conference Paper %T Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs %A Niklas Raehse %A Luregn J Schlapbach %A Daphné Chopard %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-raehse26a %I PMLR %P 1613--1666 %U https://proceedings.mlr.press/v340/raehse26a.html %V 340 %X Antimicrobial stewardship (AMS) is critical for reducing unnecessary antibiotic exposure, particularly in pediatric intensive care units (PICUs), where clinical uncertainty leads to frequent broad-spectrum use, and overuse amplifies antimicrobial resistance which have long-term consequences. Machine learning has been proposed to support AMS by identifying patient-level opportunities for intervention from electronic health record data. However, prior work focuses on adult populations and predominantly relies on static tabular representations, leaving open questions about target design, temporal modeling, and generalizability in pediatric settings. In this work, we present a systematic benchmarking study of AMS intervention prediction in the PICU. Using the publicly available Paediatric Intensive Care database from China and a private PICU cohort from the University Children’s Hospital Zurich, Switzerland, we define four clinically relevant proxy targets for reducing antibiotic use: intravenous-to-oral switching, de-escalation, discontinuation, and short-course therapy. We then compare tabular, sequence-based, and graph-based temporal models under a unified evaluation framework. We find that model performance is primarily driven by target prevalence and data characteristics rather than model complexity. Sequence models provide improvements in precision-recall trade-off over tabular approaches at coarse (24-hour) resolution, with limited additional gains when finer temporal structure is incorporated. However, these gains come at the cost of poorer calibration, with simpler tabular models producing more reliable probability estimates. Our results highlight the importance of target selection, temporal representation, and calibration in clinical machine learning, and provide practical guidance for developing reliable decision support systems for AMS in pediatric critical care.
APA
Raehse, N., Schlapbach, L.J. & Chopard, D.. (2026). Benchmarking Machine Learning Architectures for Antimicrobial Stewardship in Pediatric ICUs. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1613-1666 Available from https://proceedings.mlr.press/v340/raehse26a.html.

Related Material