CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

Anirban Porya, Mercy Prasanna Ranjit, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Sathvik Joel, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1537-1580, 2026.

Abstract

A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements on which many diagnoses depend. Today’s Vision-Language Models (VLMs) treat these as separate problems, if they address them at all — leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained with the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality — evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, VQA, and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near-parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct an off-the-shelf VLM with native tool-calling capabilities — with deterministic measurement tools while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions. We validate on rare, high-acuity ICU pathologies using clinical data from Narayana Health (NH), India, and on organ-enlargement conditions with CT-confirmed ground truth, showing that measurement-augmented CXR screening can identify high-risk cases who may require confirmatory imaging.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-porya26a, title = {CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement}, author = {Porya, Anirban and Ranjit, Mercy Prasanna and Vadlamudi, Niharika and Eathamukkala, Nikhilesh Chowdary and V, Prasanth V and Joel, Sathvik and Swamy, Abhyuday Kumara and Umredkar, Pranay Narhari and Narayan, Pradeep and Rajagopal, Vivek and Ganu, Tanuja}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1537--1580}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/porya26a/porya26a.pdf}, url = {https://proceedings.mlr.press/v340/porya26a.html}, abstract = {A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements on which many diagnoses depend. Today’s Vision-Language Models (VLMs) treat these as separate problems, if they address them at all — leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained with the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality — evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, VQA, and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near-parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct an off-the-shelf VLM with native tool-calling capabilities — with deterministic measurement tools while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions. We validate on rare, high-acuity ICU pathologies using clinical data from Narayana Health (NH), India, and on organ-enlargement conditions with CT-confirmed ground truth, showing that measurement-augmented CXR screening can identify high-risk cases who may require confirmatory imaging.} }
Endnote
%0 Conference Paper %T CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement %A Anirban Porya %A Mercy Prasanna Ranjit %A Niharika Vadlamudi %A Nikhilesh Chowdary Eathamukkala %A Prasanth V V %A Sathvik Joel %A Abhyuday Kumara Swamy %A Pranay Narhari Umredkar %A Pradeep Narayan %A Vivek Rajagopal %A Tanuja Ganu %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-porya26a %I PMLR %P 1537--1580 %U https://proceedings.mlr.press/v340/porya26a.html %V 340 %X A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements on which many diagnoses depend. Today’s Vision-Language Models (VLMs) treat these as separate problems, if they address them at all — leaving a gap between what radiologists need and what generative models provide. We introduce CARE-X, a chest X-ray VLM that narrows this gap by unifying auxiliary discriminative supervision with reward-aligned generation. CARE-X augments its generative backbone with focal-loss classification and composite-loss grounding heads, co-trained with the language-modeling objective. This auxiliary supervision produces discriminative diagnostic predictions with tunable decision thresholds and precise spatial localization while also improving report quality — evidence that structured prediction and generation reinforce one another. Building on this foundation, Decoupled Clip and Dynamic sAmpling Policy Optimization (DAPO) leverages task-specific reward signals for report generation, VQA, and spatial grounding, directly optimizing the clinical quality metrics that matter in practice. The result is state-of-the-art performance on the majority of metrics across four report generation benchmarks, 94.0% VQA accuracy on ReXVQA (+6.0 pp over the next-best baseline), and generative spatial decoding that reaches near-parity with dedicated detection heads. Separately, to address measurement-dependent diagnoses, we couple Qwen3-VL-4B-Instruct an off-the-shelf VLM with native tool-calling capabilities — with deterministic measurement tools while retaining full visual access to the image. This hybrid inference yields +43.6 pp average F1 over perception-only baselines across five measurement-dependent conditions. We validate on rare, high-acuity ICU pathologies using clinical data from Narayana Health (NH), India, and on organ-enlargement conditions with CT-confirmed ground truth, showing that measurement-augmented CXR screening can identify high-risk cases who may require confirmatory imaging.
APA
Porya, A., Ranjit, M.P., Vadlamudi, N., Eathamukkala, N.C., V, P.V., Joel, S., Swamy, A.K., Umredkar, P.N., Narayan, P., Rajagopal, V. & Ganu, T.. (2026). CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1537-1580 Available from https://proceedings.mlr.press/v340/porya26a.html.

Related Material