Towards Motion-aware Referring Image Segmentation

Chaeyun Kim, Seunghoon Yi, Yejin Kim, Yohan Jo, Joonseok Lee
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:559-567, 2026.

Abstract

Referring Image Segmentation (RIS) requires identifying objects from images based on textual descriptions. We observe that existing methods significantly underperform on motion-related queries compared to appearance-based ones, and propose to address this from both data and algorithmic perspectives. First, we introduce an efficient data augmentation scheme that extracts motion-centric phrases from original captions, exposing models to more motion expressions without additional annotations. Second, since the same object can be described differently depending on the context, we propose Multimodal Radial Contrastive Learning (MRaCL), performed on fused image-text embeddings rather than unimodal representations. For comprehensive evaluation, we introduce a new test split focusing on motion-centric queries, and introduce a new benchmark called M-Bench, where objects are distinguished primarily by actions. Extensive experiments show our method substantially improves performance on motion-centric queries across multiple RIS models, maintaining competitive results on appearance-based descriptions.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-kim26a, title = { Towards Motion-aware Referring Image Segmentation }, author = {Kim, Chaeyun and Yi, Seunghoon and Kim, Yejin and Jo, Yohan and Lee, Joonseok}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {559--567}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/kim26a/kim26a.pdf}, url = {https://proceedings.mlr.press/v300/kim26a.html}, abstract = { Referring Image Segmentation (RIS) requires identifying objects from images based on textual descriptions. We observe that existing methods significantly underperform on motion-related queries compared to appearance-based ones, and propose to address this from both data and algorithmic perspectives. First, we introduce an efficient data augmentation scheme that extracts motion-centric phrases from original captions, exposing models to more motion expressions without additional annotations. Second, since the same object can be described differently depending on the context, we propose Multimodal Radial Contrastive Learning (MRaCL), performed on fused image-text embeddings rather than unimodal representations. For comprehensive evaluation, we introduce a new test split focusing on motion-centric queries, and introduce a new benchmark called M-Bench, where objects are distinguished primarily by actions. Extensive experiments show our method substantially improves performance on motion-centric queries across multiple RIS models, maintaining competitive results on appearance-based descriptions. } }
Endnote
%0 Conference Paper %T Towards Motion-aware Referring Image Segmentation %A Chaeyun Kim %A Seunghoon Yi %A Yejin Kim %A Yohan Jo %A Joonseok Lee %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-kim26a %I PMLR %P 559--567 %U https://proceedings.mlr.press/v300/kim26a.html %V 300 %X Referring Image Segmentation (RIS) requires identifying objects from images based on textual descriptions. We observe that existing methods significantly underperform on motion-related queries compared to appearance-based ones, and propose to address this from both data and algorithmic perspectives. First, we introduce an efficient data augmentation scheme that extracts motion-centric phrases from original captions, exposing models to more motion expressions without additional annotations. Second, since the same object can be described differently depending on the context, we propose Multimodal Radial Contrastive Learning (MRaCL), performed on fused image-text embeddings rather than unimodal representations. For comprehensive evaluation, we introduce a new test split focusing on motion-centric queries, and introduce a new benchmark called M-Bench, where objects are distinguished primarily by actions. Extensive experiments show our method substantially improves performance on motion-centric queries across multiple RIS models, maintaining competitive results on appearance-based descriptions.
APA
Kim, C., Yi, S., Kim, Y., Jo, Y. & Lee, J.. (2026). Towards Motion-aware Referring Image Segmentation . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:559-567 Available from https://proceedings.mlr.press/v300/kim26a.html.

Related Material