A Cloud–Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings

Hei Ting Una Chan, Chenwei Wu, Xueshen Liu, Zesen Zhao, Boyuan Zheng, Luis Filipe Nakayama, Michael G Morley, Liyue Shen, Jiasi Chen, Z. Morley Mao
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:283-313, 2026.

Abstract

Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making requires integrating heterogeneous modalities. We introduce a cloud–edge collaborative architecture that addresses these constraints: lightweight, domain-specific models on the edge transform raw medical data into compact structured outputs, while a cloud LLM synthesizes these outputs into clinical summaries. An LLM-based orchestrator dynamically selects diagnostic tools based on patient context, promoting relevant modality coverage without processing irrelevant inputs. We evaluate on 100 multimodal clinical cases spanning cardiac, obstetric, trauma, ophthalmology, and screening scenarios — including sparse-input presentations with missing modalities and dense-input presentations with many overlapping inputs under three simulated network profiles (500 kbps–5 Mbps), reporting 95% confidence intervals throughout. The hybrid system attains the highest oracle accuracy (0.87–0.90) and the strongest factual grounding (KG precision up to 0.96), together with high coverage precision (0.95–0.99), while transmitting only $\tilde$6.5 KB of structured evidence to the cloud — three orders of magnitude less than cloud-only baselines. It maintains bandwidth-invariant latency (25–38 s) at up to 15$\times$lower token cost. These results highlight the role of architectural design in improving evidence selectivity and factual grounding, rather than merely reducing upload size, under deployment constraints.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-chan26a, title = {A Cloud–Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings}, author = {Chan, Hei Ting Una and Wu, Chenwei and Liu, Xueshen and Zhao, Zesen and Zheng, Boyuan and Nakayama, Luis Filipe and Morley, Michael G and Shen, Liyue and Chen, Jiasi and Mao, Z. Morley}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {283--313}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/chan26a/chan26a.pdf}, url = {https://proceedings.mlr.press/v340/chan26a.html}, abstract = {Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making requires integrating heterogeneous modalities. We introduce a cloud–edge collaborative architecture that addresses these constraints: lightweight, domain-specific models on the edge transform raw medical data into compact structured outputs, while a cloud LLM synthesizes these outputs into clinical summaries. An LLM-based orchestrator dynamically selects diagnostic tools based on patient context, promoting relevant modality coverage without processing irrelevant inputs. We evaluate on 100 multimodal clinical cases spanning cardiac, obstetric, trauma, ophthalmology, and screening scenarios — including sparse-input presentations with missing modalities and dense-input presentations with many overlapping inputs under three simulated network profiles (500 kbps–5 Mbps), reporting 95% confidence intervals throughout. The hybrid system attains the highest oracle accuracy (0.87–0.90) and the strongest factual grounding (KG precision up to 0.96), together with high coverage precision (0.95–0.99), while transmitting only $\tilde$6.5 KB of structured evidence to the cloud — three orders of magnitude less than cloud-only baselines. It maintains bandwidth-invariant latency (25–38 s) at up to 15$\times$lower token cost. These results highlight the role of architectural design in improving evidence selectivity and factual grounding, rather than merely reducing upload size, under deployment constraints.} }
Endnote
%0 Conference Paper %T A Cloud–Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings %A Hei Ting Una Chan %A Chenwei Wu %A Xueshen Liu %A Zesen Zhao %A Boyuan Zheng %A Luis Filipe Nakayama %A Michael G Morley %A Liyue Shen %A Jiasi Chen %A Z. Morley Mao %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-chan26a %I PMLR %P 283--313 %U https://proceedings.mlr.press/v340/chan26a.html %V 340 %X Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limited, and clinical decision-making requires integrating heterogeneous modalities. We introduce a cloud–edge collaborative architecture that addresses these constraints: lightweight, domain-specific models on the edge transform raw medical data into compact structured outputs, while a cloud LLM synthesizes these outputs into clinical summaries. An LLM-based orchestrator dynamically selects diagnostic tools based on patient context, promoting relevant modality coverage without processing irrelevant inputs. We evaluate on 100 multimodal clinical cases spanning cardiac, obstetric, trauma, ophthalmology, and screening scenarios — including sparse-input presentations with missing modalities and dense-input presentations with many overlapping inputs under three simulated network profiles (500 kbps–5 Mbps), reporting 95% confidence intervals throughout. The hybrid system attains the highest oracle accuracy (0.87–0.90) and the strongest factual grounding (KG precision up to 0.96), together with high coverage precision (0.95–0.99), while transmitting only $\tilde$6.5 KB of structured evidence to the cloud — three orders of magnitude less than cloud-only baselines. It maintains bandwidth-invariant latency (25–38 s) at up to 15$\times$lower token cost. These results highlight the role of architectural design in improving evidence selectivity and factual grounding, rather than merely reducing upload size, under deployment constraints.
APA
Chan, H.T.U., Wu, C., Liu, X., Zhao, Z., Zheng, B., Nakayama, L.F., Morley, M.G., Shen, L., Chen, J. & Mao, Z.M.. (2026). A Cloud–Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:283-313 Available from https://proceedings.mlr.press/v340/chan26a.html.

Related Material