Visual concept ranking uncovers medical shortcuts used by large multimodal models

Joseph David Janizek, Sonnet Xu, Junayd Lateef, Roxana Daneshjou
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:690-729, 2026.

Abstract

Ensuring the reliability of machine learning models in safety-critical domains such as healthcare requires auditing methods that can uncover model shortcomings. We introduce a method for identifying important visual concepts within large multimodal models (LMMs) and use it to investigate the behaviors these models exhibit when prompted with medical tasks. We primarily focus on the task of classifying malignant skin lesions from clinical dermatology images. After showing how LMMs display unexpected gaps in performance between different demographic subgroups when prompted with demonstrating examples, we apply our method, Visual Concept Ranking (VCR), to these models and prompts. VCR generates hypotheses related to different visual feature dependencies, which we are then able to validate with manual interventions.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-janizek26a, title = {Visual concept ranking uncovers medical shortcuts used by large multimodal models}, author = {Janizek, Joseph David and Xu, Sonnet and Lateef, Junayd and Daneshjou, Roxana}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {690--729}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/janizek26a/janizek26a.pdf}, url = {https://proceedings.mlr.press/v340/janizek26a.html}, abstract = {Ensuring the reliability of machine learning models in safety-critical domains such as healthcare requires auditing methods that can uncover model shortcomings. We introduce a method for identifying important visual concepts within large multimodal models (LMMs) and use it to investigate the behaviors these models exhibit when prompted with medical tasks. We primarily focus on the task of classifying malignant skin lesions from clinical dermatology images. After showing how LMMs display unexpected gaps in performance between different demographic subgroups when prompted with demonstrating examples, we apply our method, Visual Concept Ranking (VCR), to these models and prompts. VCR generates hypotheses related to different visual feature dependencies, which we are then able to validate with manual interventions.} }
Endnote
%0 Conference Paper %T Visual concept ranking uncovers medical shortcuts used by large multimodal models %A Joseph David Janizek %A Sonnet Xu %A Junayd Lateef %A Roxana Daneshjou %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-janizek26a %I PMLR %P 690--729 %U https://proceedings.mlr.press/v340/janizek26a.html %V 340 %X Ensuring the reliability of machine learning models in safety-critical domains such as healthcare requires auditing methods that can uncover model shortcomings. We introduce a method for identifying important visual concepts within large multimodal models (LMMs) and use it to investigate the behaviors these models exhibit when prompted with medical tasks. We primarily focus on the task of classifying malignant skin lesions from clinical dermatology images. After showing how LMMs display unexpected gaps in performance between different demographic subgroups when prompted with demonstrating examples, we apply our method, Visual Concept Ranking (VCR), to these models and prompts. VCR generates hypotheses related to different visual feature dependencies, which we are then able to validate with manual interventions.
APA
Janizek, J.D., Xu, S., Lateef, J. & Daneshjou, R.. (2026). Visual concept ranking uncovers medical shortcuts used by large multimodal models. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:690-729 Available from https://proceedings.mlr.press/v340/janizek26a.html.

Related Material