Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings

Utsav Dutta, Gerardo Pastrana, Sina Khoshfetrat Pakazad, Henrik Ohlsson
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:27330-27378, 2026.

Abstract

Transformer-based architectures have advanced sequence modeling in language and vision, yet general-purpose representation learning for heterogeneous multivariate time series remains underexplored. We introduce CHARM (Channel-Aware Representation Model), which incorporates channel-level textual descriptions into a Transformer encoder equivariant to channel order. CHARM is trained with a Joint Embedding Predictive Architecture (JEPA) and a novel loss promoting informative, temporally stable embeddings; latent-space prediction encourages robustness to sensor noise while description-aware gating provides interpretability through learned inter-channel relationships. Across anomaly detection, classification, and short- and long-term forecasting, the learned embeddings achieve strong performance using only a linear probe. Performance is driven primarily by the JEPA objective and conditioning architecture, with text descriptions serving as channel identifiers for cross-dataset generalization.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-dutta26b, title = {Giving Sensors a Voice: Multimodal {JEPA} for Semantic Time-Series Embeddings}, author = {Dutta, Utsav and Pastrana, Gerardo and Pakazad, Sina Khoshfetrat and Ohlsson, Henrik}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {27330--27378}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/dutta26b/dutta26b.pdf}, url = {https://proceedings.mlr.press/v306/dutta26b.html}, abstract = {Transformer-based architectures have advanced sequence modeling in language and vision, yet general-purpose representation learning for heterogeneous multivariate time series remains underexplored. We introduce CHARM (Channel-Aware Representation Model), which incorporates channel-level textual descriptions into a Transformer encoder equivariant to channel order. CHARM is trained with a Joint Embedding Predictive Architecture (JEPA) and a novel loss promoting informative, temporally stable embeddings; latent-space prediction encourages robustness to sensor noise while description-aware gating provides interpretability through learned inter-channel relationships. Across anomaly detection, classification, and short- and long-term forecasting, the learned embeddings achieve strong performance using only a linear probe. Performance is driven primarily by the JEPA objective and conditioning architecture, with text descriptions serving as channel identifiers for cross-dataset generalization.} }
Endnote
%0 Conference Paper %T Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings %A Utsav Dutta %A Gerardo Pastrana %A Sina Khoshfetrat Pakazad %A Henrik Ohlsson %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-dutta26b %I PMLR %P 27330--27378 %U https://proceedings.mlr.press/v306/dutta26b.html %V 306 %X Transformer-based architectures have advanced sequence modeling in language and vision, yet general-purpose representation learning for heterogeneous multivariate time series remains underexplored. We introduce CHARM (Channel-Aware Representation Model), which incorporates channel-level textual descriptions into a Transformer encoder equivariant to channel order. CHARM is trained with a Joint Embedding Predictive Architecture (JEPA) and a novel loss promoting informative, temporally stable embeddings; latent-space prediction encourages robustness to sensor noise while description-aware gating provides interpretability through learned inter-channel relationships. Across anomaly detection, classification, and short- and long-term forecasting, the learned embeddings achieve strong performance using only a linear probe. Performance is driven primarily by the JEPA objective and conditioning architecture, with text descriptions serving as channel identifiers for cross-dataset generalization.
APA
Dutta, U., Pastrana, G., Pakazad, S.K. & Ohlsson, H.. (2026). Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:27330-27378 Available from https://proceedings.mlr.press/v306/dutta26b.html.

Related Material