FocusViT: Faithful Explanations for Vision Transformers via Gradient-Guided Layer-Skipping

Mohsin Ali, Haider Raza, John Q Gan, Muhammad Haris Khan
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:1522-1530, 2026.

Abstract

Vision Transformers (ViTs) have emerged as powerful alternatives to CNNs for various vision tasks, yet their token-based, attention-driven architecture makes interpreting their predictions challenging. Existing explainability methods, such as Grad-CAM and Attention Rollout, either fail to capture hierarchical semantic information or assume attention directly reflects importance, often leading to misleading explanations. We propose FocusViT, a novel explainability framework that integrates gradient-weighted attention attribution with validation-based, faithfulness-driven layer aggregation. By fusing attention maps with class-specific gradients and introducing per-head dynamic weighting, FocusViT highlights not only where the model attends but also how sensitive the prediction is to those attentions. Furthermore, our adaptive layer-skipping strategy ensures that only semantically meaningful layers contribute to the final explanation, enhancing both faithfulness and clarity. Extensive quantitative and qualitative evaluations on diverse benchmarks demonstrate that FocusViT improves over existing methods in faithfulness and sparsity, achieving competitive robustness and class sensitivity, and provides sharper, more reliable visual explanations for ViTs. The official implementation is publicly available at: \url{https://github.com/game-sys/focusvit-aistats2026.git}

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-ali26a, title = { FocusViT: Faithful Explanations for Vision Transformers via Gradient-Guided Layer-Skipping }, author = {Ali, Mohsin and Raza, Haider and Gan, John Q and Khan, Muhammad Haris}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {1522--1530}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/ali26a/ali26a.pdf}, url = {https://proceedings.mlr.press/v300/ali26a.html}, abstract = { Vision Transformers (ViTs) have emerged as powerful alternatives to CNNs for various vision tasks, yet their token-based, attention-driven architecture makes interpreting their predictions challenging. Existing explainability methods, such as Grad-CAM and Attention Rollout, either fail to capture hierarchical semantic information or assume attention directly reflects importance, often leading to misleading explanations. We propose FocusViT, a novel explainability framework that integrates gradient-weighted attention attribution with validation-based, faithfulness-driven layer aggregation. By fusing attention maps with class-specific gradients and introducing per-head dynamic weighting, FocusViT highlights not only where the model attends but also how sensitive the prediction is to those attentions. Furthermore, our adaptive layer-skipping strategy ensures that only semantically meaningful layers contribute to the final explanation, enhancing both faithfulness and clarity. Extensive quantitative and qualitative evaluations on diverse benchmarks demonstrate that FocusViT improves over existing methods in faithfulness and sparsity, achieving competitive robustness and class sensitivity, and provides sharper, more reliable visual explanations for ViTs. The official implementation is publicly available at: \url{https://github.com/game-sys/focusvit-aistats2026.git} } }
Endnote
%0 Conference Paper %T FocusViT: Faithful Explanations for Vision Transformers via Gradient-Guided Layer-Skipping %A Mohsin Ali %A Haider Raza %A John Q Gan %A Muhammad Haris Khan %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-ali26a %I PMLR %P 1522--1530 %U https://proceedings.mlr.press/v300/ali26a.html %V 300 %X Vision Transformers (ViTs) have emerged as powerful alternatives to CNNs for various vision tasks, yet their token-based, attention-driven architecture makes interpreting their predictions challenging. Existing explainability methods, such as Grad-CAM and Attention Rollout, either fail to capture hierarchical semantic information or assume attention directly reflects importance, often leading to misleading explanations. We propose FocusViT, a novel explainability framework that integrates gradient-weighted attention attribution with validation-based, faithfulness-driven layer aggregation. By fusing attention maps with class-specific gradients and introducing per-head dynamic weighting, FocusViT highlights not only where the model attends but also how sensitive the prediction is to those attentions. Furthermore, our adaptive layer-skipping strategy ensures that only semantically meaningful layers contribute to the final explanation, enhancing both faithfulness and clarity. Extensive quantitative and qualitative evaluations on diverse benchmarks demonstrate that FocusViT improves over existing methods in faithfulness and sparsity, achieving competitive robustness and class sensitivity, and provides sharper, more reliable visual explanations for ViTs. The official implementation is publicly available at: \url{https://github.com/game-sys/focusvit-aistats2026.git}
APA
Ali, M., Raza, H., Gan, J.Q. & Khan, M.H.. (2026). FocusViT: Faithful Explanations for Vision Transformers via Gradient-Guided Layer-Skipping . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:1522-1530 Available from https://proceedings.mlr.press/v300/ali26a.html.

Related Material