Human–AI Deferral under Limited Expert Availability: A Study in Intraoperative Ischemia Detection

Nihal Murali, Karim Elzokm, Parthasarathy Thirumala, kayhan Batmanghelich, Shyam Visweswaran
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1367-1402, 2026.

Abstract

Intraoperative neuromonitoring (IONM) during carotid endarterectomy (CEA) is used to monitor cerebral ischemia, but reliable interpretation depends on scarce expert neurophysiologists. Because many medical tasks are safety-critical, relying solely on artificial intelligence (AI) is impractical. We study how human–AI deferral methods perform in this setting and also introduce Learning to Augment (L2A), a simple method that selectively incorporates human input into AI predictions. We evaluate these methods on intraoperative electroencephalographic data from 400 CEA cases in a large U.S. academic health system. Unlike prior deferral work that focuses on non-clinical tasks and relies on expert humans, we use novice monitors with no prior IONM experience and only brief training. We use these novices as a conservative lower-bound test of whether non-expert human input contains complementary information, not as a proposed replacement for clinically trained personnel. Across methods, human–AI deferral shows that substantial gains over the AI alone can be achieved with minimal human involvement, even when the human is a novice. With just 5–10% novice involvement, L2A improves precision and sensitivity by $\sim$40%, and the area under the precision-recall curve (AUPRC) by $\sim$20% compared to the AI model alone. Importantly, while some deferral methods degrade with less experienced human input, others such as L2A remain robust with better calibration. We further observe that these methods produce sparse and temporally clustered deferral decisions, enabling long periods without human intervention. These properties make human–AI deferral approaches practical in settings where expert neurophysiologists are not available.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-murali26a, title = {Human–AI Deferral under Limited Expert Availability: A Study in Intraoperative Ischemia Detection}, author = {Murali, Nihal and Elzokm, Karim and Thirumala, Parthasarathy and Batmanghelich, kayhan and Visweswaran, Shyam}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1367--1402}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/murali26a/murali26a.pdf}, url = {https://proceedings.mlr.press/v340/murali26a.html}, abstract = {Intraoperative neuromonitoring (IONM) during carotid endarterectomy (CEA) is used to monitor cerebral ischemia, but reliable interpretation depends on scarce expert neurophysiologists. Because many medical tasks are safety-critical, relying solely on artificial intelligence (AI) is impractical. We study how human–AI deferral methods perform in this setting and also introduce Learning to Augment (L2A), a simple method that selectively incorporates human input into AI predictions. We evaluate these methods on intraoperative electroencephalographic data from 400 CEA cases in a large U.S. academic health system. Unlike prior deferral work that focuses on non-clinical tasks and relies on expert humans, we use novice monitors with no prior IONM experience and only brief training. We use these novices as a conservative lower-bound test of whether non-expert human input contains complementary information, not as a proposed replacement for clinically trained personnel. Across methods, human–AI deferral shows that substantial gains over the AI alone can be achieved with minimal human involvement, even when the human is a novice. With just 5–10% novice involvement, L2A improves precision and sensitivity by $\sim$40%, and the area under the precision-recall curve (AUPRC) by $\sim$20% compared to the AI model alone. Importantly, while some deferral methods degrade with less experienced human input, others such as L2A remain robust with better calibration. We further observe that these methods produce sparse and temporally clustered deferral decisions, enabling long periods without human intervention. These properties make human–AI deferral approaches practical in settings where expert neurophysiologists are not available.} }
Endnote
%0 Conference Paper %T Human–AI Deferral under Limited Expert Availability: A Study in Intraoperative Ischemia Detection %A Nihal Murali %A Karim Elzokm %A Parthasarathy Thirumala %A kayhan Batmanghelich %A Shyam Visweswaran %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-murali26a %I PMLR %P 1367--1402 %U https://proceedings.mlr.press/v340/murali26a.html %V 340 %X Intraoperative neuromonitoring (IONM) during carotid endarterectomy (CEA) is used to monitor cerebral ischemia, but reliable interpretation depends on scarce expert neurophysiologists. Because many medical tasks are safety-critical, relying solely on artificial intelligence (AI) is impractical. We study how human–AI deferral methods perform in this setting and also introduce Learning to Augment (L2A), a simple method that selectively incorporates human input into AI predictions. We evaluate these methods on intraoperative electroencephalographic data from 400 CEA cases in a large U.S. academic health system. Unlike prior deferral work that focuses on non-clinical tasks and relies on expert humans, we use novice monitors with no prior IONM experience and only brief training. We use these novices as a conservative lower-bound test of whether non-expert human input contains complementary information, not as a proposed replacement for clinically trained personnel. Across methods, human–AI deferral shows that substantial gains over the AI alone can be achieved with minimal human involvement, even when the human is a novice. With just 5–10% novice involvement, L2A improves precision and sensitivity by $\sim$40%, and the area under the precision-recall curve (AUPRC) by $\sim$20% compared to the AI model alone. Importantly, while some deferral methods degrade with less experienced human input, others such as L2A remain robust with better calibration. We further observe that these methods produce sparse and temporally clustered deferral decisions, enabling long periods without human intervention. These properties make human–AI deferral approaches practical in settings where expert neurophysiologists are not available.
APA
Murali, N., Elzokm, K., Thirumala, P., Batmanghelich, k. & Visweswaran, S.. (2026). Human–AI Deferral under Limited Expert Availability: A Study in Intraoperative Ischemia Detection. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1367-1402 Available from https://proceedings.mlr.press/v340/murali26a.html.

Related Material