Enhancing LLMs’ Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry

Junu Kim, Chaeeun Shim, Sungjin Park, Su Yeon Lee, Gee Young Suh, Seong Jin Choi, Song Mi Moon, Kyoung-Ho Song, Eu Suk Kim, Hong Bin Kim, Sejoong Kim, Chami Im, Dong-Wan Kang, Yong Soo Kim, Hee-Joon Bae, Sung Yoon Lim, Han-Gil Jeong, Edward Choi
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:835-885, 2026.

Abstract

Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited. This limitation is likely due to insufficient exposure to real-world clinical data during training, as such data are typically excluded because of privacy concerns. To address this gap, we trained an LLM on real-world clinical data from a nationwide sepsis registry and evaluated the reasoning improvements across diverse datasets and tasks. The trained model demonstrated strong clinical reasoning performance on in-domain test sets, supported by both quantitative metrics and expert evaluations. Moreover, these enhanced reasoning capabilities generalized to an external sepsis dataset involving different tasks and patient cohorts, an open-ended antibiotic consultation task, and a disease beyond sepsis. Future research should focus on training LLMs on large-scale, multi-disease clinical datasets to enable more powerful and general-purpose clinical reasoning models.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-kim26a, title = {Enhancing LLMs’ Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry}, author = {Kim, Junu and Shim, Chaeeun and Park, Sungjin and Lee, Su Yeon and Suh, Gee Young and Choi, Seong Jin and Moon, Song Mi and Song, Kyoung-Ho and Kim, Eu Suk and Kim, Hong Bin and Kim, Sejoong and Im, Chami and Kang, Dong-Wan and Kim, Yong Soo and Bae, Hee-Joon and Lim, Sung Yoon and Jeong, Han-Gil and Choi, Edward}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {835--885}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/kim26a/kim26a.pdf}, url = {https://proceedings.mlr.press/v340/kim26a.html}, abstract = {Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited. This limitation is likely due to insufficient exposure to real-world clinical data during training, as such data are typically excluded because of privacy concerns. To address this gap, we trained an LLM on real-world clinical data from a nationwide sepsis registry and evaluated the reasoning improvements across diverse datasets and tasks. The trained model demonstrated strong clinical reasoning performance on in-domain test sets, supported by both quantitative metrics and expert evaluations. Moreover, these enhanced reasoning capabilities generalized to an external sepsis dataset involving different tasks and patient cohorts, an open-ended antibiotic consultation task, and a disease beyond sepsis. Future research should focus on training LLMs on large-scale, multi-disease clinical datasets to enable more powerful and general-purpose clinical reasoning models.} }
Endnote
%0 Conference Paper %T Enhancing LLMs’ Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry %A Junu Kim %A Chaeeun Shim %A Sungjin Park %A Su Yeon Lee %A Gee Young Suh %A Seong Jin Choi %A Song Mi Moon %A Kyoung-Ho Song %A Eu Suk Kim %A Hong Bin Kim %A Sejoong Kim %A Chami Im %A Dong-Wan Kang %A Yong Soo Kim %A Hee-Joon Bae %A Sung Yoon Lim %A Han-Gil Jeong %A Edward Choi %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-kim26a %I PMLR %P 835--885 %U https://proceedings.mlr.press/v340/kim26a.html %V 340 %X Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in real-world clinical practice remains limited. This limitation is likely due to insufficient exposure to real-world clinical data during training, as such data are typically excluded because of privacy concerns. To address this gap, we trained an LLM on real-world clinical data from a nationwide sepsis registry and evaluated the reasoning improvements across diverse datasets and tasks. The trained model demonstrated strong clinical reasoning performance on in-domain test sets, supported by both quantitative metrics and expert evaluations. Moreover, these enhanced reasoning capabilities generalized to an external sepsis dataset involving different tasks and patient cohorts, an open-ended antibiotic consultation task, and a disease beyond sepsis. Future research should focus on training LLMs on large-scale, multi-disease clinical datasets to enable more powerful and general-purpose clinical reasoning models.
APA
Kim, J., Shim, C., Park, S., Lee, S.Y., Suh, G.Y., Choi, S.J., Moon, S.M., Song, K., Kim, E.S., Kim, H.B., Kim, S., Im, C., Kang, D., Kim, Y.S., Bae, H., Lim, S.Y., Jeong, H. & Choi, E.. (2026). Enhancing LLMs’ Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:835-885 Available from https://proceedings.mlr.press/v340/kim26a.html.

Related Material