Characterizing Population Gaps in Clinical Decision-Making: A Guideline-Based Benchmark and Population-Aware Retrieval Analysis in Anesthesiology

Chaiho Shin, Jung-Bin Park, Kwangsoo Kim, Hee-Soo Kim
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1791-1833, 2026.

Abstract

Large language models (LLMs) have shown strong performance on medical question answering, yet their ability to support population-specific, action-oriented clinical decision-making remains underexplored. Our analysis of five widely used medical QA benchmarks shows that pediatric anesthesia accounts for only 0.2% of all questions, revealing a critical evaluation gap. To address this, we construct PopAnesQA, a population-aware anesthesiology question-answering benchmark grounded in clinical guidelines, and introduce Prof-RAG, a diagnostic framework for isolating the effects of patient-specific context on retrieval and downstream reasoning. Across diverse LLMs, we observe a consistent pediatric performance gap relative to the general subset, with differences exceeding 15 percentage points even in proprietary models like GPT-4o. We find that incorporating patient profile information leads to highly model-dependent effects, improving performance in some models while degrading it in others. Retrieval trajectory analysis shows that incorporating profile information increases the proportion of population-aligned documents, but does not consistently improve relevance to the clinical question. Overall, our findings highlight that effective clinical retrieval requires not only population awareness, but also careful alignment between patient-specific context and clinical intent.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-shin26a, title = {Characterizing Population Gaps in Clinical Decision-Making: A Guideline-Based Benchmark and Population-Aware Retrieval Analysis in Anesthesiology}, author = {Shin, Chaiho and Park, Jung-Bin and Kim, Kwangsoo and Kim, Hee-Soo}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1791--1833}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/shin26a/shin26a.pdf}, url = {https://proceedings.mlr.press/v340/shin26a.html}, abstract = {Large language models (LLMs) have shown strong performance on medical question answering, yet their ability to support population-specific, action-oriented clinical decision-making remains underexplored. Our analysis of five widely used medical QA benchmarks shows that pediatric anesthesia accounts for only 0.2% of all questions, revealing a critical evaluation gap. To address this, we construct PopAnesQA, a population-aware anesthesiology question-answering benchmark grounded in clinical guidelines, and introduce Prof-RAG, a diagnostic framework for isolating the effects of patient-specific context on retrieval and downstream reasoning. Across diverse LLMs, we observe a consistent pediatric performance gap relative to the general subset, with differences exceeding 15 percentage points even in proprietary models like GPT-4o. We find that incorporating patient profile information leads to highly model-dependent effects, improving performance in some models while degrading it in others. Retrieval trajectory analysis shows that incorporating profile information increases the proportion of population-aligned documents, but does not consistently improve relevance to the clinical question. Overall, our findings highlight that effective clinical retrieval requires not only population awareness, but also careful alignment between patient-specific context and clinical intent.} }
Endnote
%0 Conference Paper %T Characterizing Population Gaps in Clinical Decision-Making: A Guideline-Based Benchmark and Population-Aware Retrieval Analysis in Anesthesiology %A Chaiho Shin %A Jung-Bin Park %A Kwangsoo Kim %A Hee-Soo Kim %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-shin26a %I PMLR %P 1791--1833 %U https://proceedings.mlr.press/v340/shin26a.html %V 340 %X Large language models (LLMs) have shown strong performance on medical question answering, yet their ability to support population-specific, action-oriented clinical decision-making remains underexplored. Our analysis of five widely used medical QA benchmarks shows that pediatric anesthesia accounts for only 0.2% of all questions, revealing a critical evaluation gap. To address this, we construct PopAnesQA, a population-aware anesthesiology question-answering benchmark grounded in clinical guidelines, and introduce Prof-RAG, a diagnostic framework for isolating the effects of patient-specific context on retrieval and downstream reasoning. Across diverse LLMs, we observe a consistent pediatric performance gap relative to the general subset, with differences exceeding 15 percentage points even in proprietary models like GPT-4o. We find that incorporating patient profile information leads to highly model-dependent effects, improving performance in some models while degrading it in others. Retrieval trajectory analysis shows that incorporating profile information increases the proportion of population-aligned documents, but does not consistently improve relevance to the clinical question. Overall, our findings highlight that effective clinical retrieval requires not only population awareness, but also careful alignment between patient-specific context and clinical intent.
APA
Shin, C., Park, J., Kim, K. & Kim, H.. (2026). Characterizing Population Gaps in Clinical Decision-Making: A Guideline-Based Benchmark and Population-Aware Retrieval Analysis in Anesthesiology. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1791-1833 Available from https://proceedings.mlr.press/v340/shin26a.html.

Related Material