The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space

Bing-Cheng Chuang, I-Hsuan Chu, Bor-Jiun Lin, Yuanfu Yang, Min Sun, Chun-Yi Lee
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:20776-20798, 2026.

Abstract

Diffusion-based Vision-Language-Action policies achieve remarkable success in robotic manipulation, yet commit a fundamental geometric error we term the Euclidean Fallacy: representing SE(3) poses as flat $\mathbb{R}^{12}$ vectors. This approximation induces (1) manifold drift violating SO(3) constraints, (2) broken equivariance under coordinate transformations, and (3) non-geodesic trajectories with excessive kinematic cost. We introduce Lie Diffuser Actor (LDA), a diffusion framework operating intrinsically on SE(3). Our method injects noise through left-invariant SDEs, predicts scores in the tangent space, and retracts samples via the exponential map. This formulation eliminates manifold drift by construction while guaranteeing coordinate-frame equivariance and geodesic optimality. On CALVIN ABC$\rightarrow$D, LDA improves average task length from $3.27$ to $3.51$ ($+7.3%$). We further validate our method on real robot and the results show that our methodology outperforms the baseline on majority tasks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chuang26a, title = {The Lie We Tell: Correcting the {E}uclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space}, author = {Chuang, Bing-Cheng and Chu, I-Hsuan and Lin, Bor-Jiun and Yang, Yuanfu and Sun, Min and Lee, Chun-Yi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {20776--20798}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chuang26a/chuang26a.pdf}, url = {https://proceedings.mlr.press/v306/chuang26a.html}, abstract = {Diffusion-based Vision-Language-Action policies achieve remarkable success in robotic manipulation, yet commit a fundamental geometric error we term the Euclidean Fallacy: representing SE(3) poses as flat $\mathbb{R}^{12}$ vectors. This approximation induces (1) manifold drift violating SO(3) constraints, (2) broken equivariance under coordinate transformations, and (3) non-geodesic trajectories with excessive kinematic cost. We introduce Lie Diffuser Actor (LDA), a diffusion framework operating intrinsically on SE(3). Our method injects noise through left-invariant SDEs, predicts scores in the tangent space, and retracts samples via the exponential map. This formulation eliminates manifold drift by construction while guaranteeing coordinate-frame equivariance and geodesic optimality. On CALVIN ABC$\rightarrow$D, LDA improves average task length from $3.27$ to $3.51$ ($+7.3%$). We further validate our method on real robot and the results show that our methodology outperforms the baseline on majority tasks.} }
Endnote
%0 Conference Paper %T The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space %A Bing-Cheng Chuang %A I-Hsuan Chu %A Bor-Jiun Lin %A Yuanfu Yang %A Min Sun %A Chun-Yi Lee %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chuang26a %I PMLR %P 20776--20798 %U https://proceedings.mlr.press/v306/chuang26a.html %V 306 %X Diffusion-based Vision-Language-Action policies achieve remarkable success in robotic manipulation, yet commit a fundamental geometric error we term the Euclidean Fallacy: representing SE(3) poses as flat $\mathbb{R}^{12}$ vectors. This approximation induces (1) manifold drift violating SO(3) constraints, (2) broken equivariance under coordinate transformations, and (3) non-geodesic trajectories with excessive kinematic cost. We introduce Lie Diffuser Actor (LDA), a diffusion framework operating intrinsically on SE(3). Our method injects noise through left-invariant SDEs, predicts scores in the tangent space, and retracts samples via the exponential map. This formulation eliminates manifold drift by construction while guaranteeing coordinate-frame equivariance and geodesic optimality. On CALVIN ABC$\rightarrow$D, LDA improves average task length from $3.27$ to $3.51$ ($+7.3%$). We further validate our method on real robot and the results show that our methodology outperforms the baseline on majority tasks.
APA
Chuang, B., Chu, I., Lin, B., Yang, Y., Sun, M. & Lee, C.. (2026). The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:20776-20798 Available from https://proceedings.mlr.press/v306/chuang26a.html.

Related Material