Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration

Dongkyu Cho, Miao Zhang, Gregory D Lyng, Rumi Chunara
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:19904-19921, 2026.

Abstract

Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Compared to existing LLM-based and traditional augmentation methods, our generated data significantly improves preservation of critical medical information and reduces hallucinations at both the token and concept levels. Experiments on downstream clinical prediction tasks demonstrate consistent performance gains over existing augmentation methods. This lightweight collaborative framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-cho26i, title = {Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration}, author = {Cho, Dongkyu and Zhang, Miao and Lyng, Gregory D and Chunara, Rumi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {19904--19921}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/cho26i/cho26i.pdf}, url = {https://proceedings.mlr.press/v306/cho26i.html}, abstract = {Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Compared to existing LLM-based and traditional augmentation methods, our generated data significantly improves preservation of critical medical information and reduces hallucinations at both the token and concept levels. Experiments on downstream clinical prediction tasks demonstrate consistent performance gains over existing augmentation methods. This lightweight collaborative framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.} }
Endnote
%0 Conference Paper %T Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration %A Dongkyu Cho %A Miao Zhang %A Gregory D Lyng %A Rumi Chunara %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-cho26i %I PMLR %P 19904--19921 %U https://proceedings.mlr.press/v306/cho26i.html %V 306 %X Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Compared to existing LLM-based and traditional augmentation methods, our generated data significantly improves preservation of critical medical information and reduces hallucinations at both the token and concept levels. Experiments on downstream clinical prediction tasks demonstrate consistent performance gains over existing augmentation methods. This lightweight collaborative framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.
APA
Cho, D., Zhang, M., Lyng, G.D. & Chunara, R.. (2026). Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:19904-19921 Available from https://proceedings.mlr.press/v306/cho26i.html.

Related Material