[edit]
Enhancing LLMs’ Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:835-885, 2026.
Abstract
Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited. This limitation is likely due to insufficient exposure to real-world clinical data during training, as such data are typically excluded because of privacy concerns. To address this gap, we trained an LLM on real-world clinical data from a nationwide sepsis registry and evaluated the reasoning improvements across diverse datasets and tasks. The trained model demonstrated strong clinical reasoning performance on in-domain test sets, supported by both quantitative metrics and expert evaluations. Moreover, these enhanced reasoning capabilities generalized to an external sepsis dataset involving different tasks and patient cohorts, an open-ended antibiotic consultation task, and a disease beyond sepsis. Future research should focus on training LLMs on large-scale, multi-disease clinical datasets to enable more powerful and general-purpose clinical reasoning models.