RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction

Yuwei Zhang; Tong Xia; Aaqib Saeed; Cecilia Mascolo

RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction

Yuwei Zhang, Tong Xia, Aaqib Saeed, Cecilia Mascolo

Proceedings of the 4th Machine Learning for Health Symposium, PMLR 259:1053-1066, 2025.

Abstract

The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific design. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.

Cite this Paper

BibTeX

@InProceedings{pmlr-v259-zhang25a,
  title = 	 {RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction},
  author =       {Zhang, Yuwei and Xia, Tong and Saeed, Aaqib and Mascolo, Cecilia},
  booktitle = 	 {Proceedings of the 4th Machine Learning for Health Symposium},
  pages = 	 {1053--1066},
  year = 	 {2025},
  editor = 	 {Hegselmann, Stefan and Zhou, Helen and Healey, Elizabeth and Chang, Trenton and Ellington, Caleb and Mhasawade, Vishwali and Tonekaboni, Sana and Argaw, Peniel and Zhang, Haoran},
  volume = 	 {259},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {15--16 Dec},
  publisher =    {PMLR},
  pdf = 	 {https://raw.githubusercontent.com/mlresearch/v259/main/assets/zhang25a/zhang25a.pdf},
  url = 	 {https://proceedings.mlr.press/v259/zhang25a.html},
  abstract = 	 {The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific design. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.}
}

Endnote

%0 Conference Paper
%T RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
%A Yuwei Zhang
%A Tong Xia
%A Aaqib Saeed
%A Cecilia Mascolo
%B Proceedings of the 4th Machine Learning for Health Symposium
%C Proceedings of Machine Learning Research
%D 2025
%E Stefan Hegselmann
%E Helen Zhou
%E Elizabeth Healey
%E Trenton Chang
%E Caleb Ellington
%E Vishwali Mhasawade
%E Sana Tonekaboni
%E Peniel Argaw
%E Haoran Zhang	
%F pmlr-v259-zhang25a
%I PMLR
%P 1053--1066
%U https://proceedings.mlr.press/v259/zhang25a.html
%V 259
%X The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific design. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.

APA

Zhang, Y., Xia, T., Saeed, A. & Mascolo, C.. (2025). RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction. Proceedings of the 4th Machine Learning for Health Symposium, in Proceedings of Machine Learning Research 259:1053-1066 Available from https://proceedings.mlr.press/v259/zhang25a.html.

RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction

Abstract

Cite this Paper

Related Material