MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction

Meng Chen, Junjie Yang, Zechen Li, Kai Zhao, Hongjun Dai, Weiming Huang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:17800-17814, 2026.

Abstract

Predicting street-level socioeconomic indicators from street view imagery is fundamental to urban planning. Existing methods typically extract visual features via pretrained encoders and propagate information through graph-based learning, but they fail to fully exploit the structured, task-relevant, and label-efficient learning signals inherent in urban scenes. We propose MetaStreet, a semi-supervised multimodal framework with three components: (1) a semantic-spatial visual encoder that jointly models object co-occurrence and spatial adjacency at the semantic category level, (2) a task-aware textual encoder that steers LLMs toward prediction-relevant features via task-specific prompts, and (3) a geography-aware graph contrastive learning module that leverages spatial autocorrelation to extend contrastive supervision to unlabeled streets, enabling them to actively participate in representation learning. Experiments on two cities across three socioeconomic prediction tasks demonstrate that MetaStreet consistently outperforms state-of-the-art methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26fo, title = {{M}eta{S}treet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction}, author = {Chen, Meng and Yang, Junjie and Li, Zechen and Zhao, Kai and Dai, Hongjun and Huang, Weiming}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {17800--17814}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26fo/chen26fo.pdf}, url = {https://proceedings.mlr.press/v306/chen26fo.html}, abstract = {Predicting street-level socioeconomic indicators from street view imagery is fundamental to urban planning. Existing methods typically extract visual features via pretrained encoders and propagate information through graph-based learning, but they fail to fully exploit the structured, task-relevant, and label-efficient learning signals inherent in urban scenes. We propose MetaStreet, a semi-supervised multimodal framework with three components: (1) a semantic-spatial visual encoder that jointly models object co-occurrence and spatial adjacency at the semantic category level, (2) a task-aware textual encoder that steers LLMs toward prediction-relevant features via task-specific prompts, and (3) a geography-aware graph contrastive learning module that leverages spatial autocorrelation to extend contrastive supervision to unlabeled streets, enabling them to actively participate in representation learning. Experiments on two cities across three socioeconomic prediction tasks demonstrate that MetaStreet consistently outperforms state-of-the-art methods.} }
Endnote
%0 Conference Paper %T MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction %A Meng Chen %A Junjie Yang %A Zechen Li %A Kai Zhao %A Hongjun Dai %A Weiming Huang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26fo %I PMLR %P 17800--17814 %U https://proceedings.mlr.press/v306/chen26fo.html %V 306 %X Predicting street-level socioeconomic indicators from street view imagery is fundamental to urban planning. Existing methods typically extract visual features via pretrained encoders and propagate information through graph-based learning, but they fail to fully exploit the structured, task-relevant, and label-efficient learning signals inherent in urban scenes. We propose MetaStreet, a semi-supervised multimodal framework with three components: (1) a semantic-spatial visual encoder that jointly models object co-occurrence and spatial adjacency at the semantic category level, (2) a task-aware textual encoder that steers LLMs toward prediction-relevant features via task-specific prompts, and (3) a geography-aware graph contrastive learning module that leverages spatial autocorrelation to extend contrastive supervision to unlabeled streets, enabling them to actively participate in representation learning. Experiments on two cities across three socioeconomic prediction tasks demonstrate that MetaStreet consistently outperforms state-of-the-art methods.
APA
Chen, M., Yang, J., Li, Z., Zhao, K., Dai, H. & Huang, W.. (2026). MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:17800-17814 Available from https://proceedings.mlr.press/v306/chen26fo.html.

Related Material