[edit]
Characterizing Population Gaps in Clinical Decision-Making: A Guideline-Based Benchmark and Population-Aware Retrieval Analysis in Anesthesiology
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1791-1833, 2026.
Abstract
Large language models (LLMs) have shown strong performance on medical question answering, yet their ability to support population-specific, action-oriented clinical decision-making remains underexplored. Our analysis of five widely used medical QA benchmarks shows that pediatric anesthesia accounts for only 0.2% of all questions, revealing a critical evaluation gap. To address this, we construct PopAnesQA, a population-aware anesthesiology question-answering benchmark grounded in clinical guidelines, and introduce Prof-RAG, a diagnostic framework for isolating the effects of patient-specific context on retrieval and downstream reasoning. Across diverse LLMs, we observe a consistent pediatric performance gap relative to the general subset, with differences exceeding 15 percentage points even in proprietary models like GPT-4o. We find that incorporating patient profile information leads to highly model-dependent effects, improving performance in some models while degrading it in others. Retrieval trajectory analysis shows that incorporating profile information increases the proportion of population-aligned documents, but does not consistently improve relevance to the clinical question. Overall, our findings highlight that effective clinical retrieval requires not only population awareness, but also careful alignment between patient-specific context and clinical intent.