CLERF: Contrastive LEaRning for Full-Range Head Pose Estimation

Ting-Ruen Wei, Huei-Chung Hu, Haowei Liu, Xuyang Wu, Yi Fang, Hsin-Tai Wu
Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling, PMLR 326:519-532, 2026.

Abstract

We propose a novel framework for representation learning in head pose estimation (HPE) that overcomes the challenges posed by sparse head pose data, which previously made triplet sampling infeasible. Leveraging recent advances in 3D-aware generative adversarial networks (3D GANs), we generate anchor–positive–negative triplets and perform contrastive learning on extensively augmented data, including geometric transformations. This enables the network to learn robust, geometry-aware representations that improve HPE accuracy. We observe that existing HPE models struggle when test images are slightly rotated or flipped, while our method maintains strong performance. Experiments show that our framework matches state-of-the-art models on standard test sets and outperforms them on augmented and full-range poses. Our model handles full-range HPE, accurately predicting head poses across the entire rotation spectrum, including upside-down orientations, and outperforms existing full-yaw range methods.

Cite this Paper


BibTeX
@InProceedings{pmlr-v326-wei26a, title = {CLERF: Contrastive LEaRning for Full-Range Head Pose Estimation}, author = {Wei, Ting-Ruen and Hu, Huei-Chung and Liu, Haowei and Wu, Xuyang and Fang, Yi and Wu, Hsin-Tai}, booktitle = {Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling}, pages = {519--532}, year = {2026}, editor = {Pouplin, Alison and Vadgama, Sharvaree and Bekkers, Erik and Kaba, Sékou-Oumar and Lawrence, Hannah and Lecha, Manuel and Baker, Elizabeth and Suk, Julian and Walters, Robin and Tomczak, Jakub and Jegelka, Stefanie}, volume = {326}, series = {Proceedings of Machine Learning Research}, month = {26 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v326/main/assets/wei26a/wei26a.pdf}, url = {https://proceedings.mlr.press/v326/wei26a.html}, abstract = {We propose a novel framework for representation learning in head pose estimation (HPE) that overcomes the challenges posed by sparse head pose data, which previously made triplet sampling infeasible. Leveraging recent advances in 3D-aware generative adversarial networks (3D GANs), we generate anchor–positive–negative triplets and perform contrastive learning on extensively augmented data, including geometric transformations. This enables the network to learn robust, geometry-aware representations that improve HPE accuracy. We observe that existing HPE models struggle when test images are slightly rotated or flipped, while our method maintains strong performance. Experiments show that our framework matches state-of-the-art models on standard test sets and outperforms them on augmented and full-range poses. Our model handles full-range HPE, accurately predicting head poses across the entire rotation spectrum, including upside-down orientations, and outperforms existing full-yaw range methods.} }
Endnote
%0 Conference Paper %T CLERF: Contrastive LEaRning for Full-Range Head Pose Estimation %A Ting-Ruen Wei %A Huei-Chung Hu %A Haowei Liu %A Xuyang Wu %A Yi Fang %A Hsin-Tai Wu %B Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling %C Proceedings of Machine Learning Research %D 2026 %E Alison Pouplin %E Sharvaree Vadgama %E Erik Bekkers %E Sékou-Oumar Kaba %E Hannah Lawrence %E Manuel Lecha %E Elizabeth Baker %E Julian Suk %E Robin Walters %E Jakub Tomczak %E Stefanie Jegelka %F pmlr-v326-wei26a %I PMLR %P 519--532 %U https://proceedings.mlr.press/v326/wei26a.html %V 326 %X We propose a novel framework for representation learning in head pose estimation (HPE) that overcomes the challenges posed by sparse head pose data, which previously made triplet sampling infeasible. Leveraging recent advances in 3D-aware generative adversarial networks (3D GANs), we generate anchor–positive–negative triplets and perform contrastive learning on extensively augmented data, including geometric transformations. This enables the network to learn robust, geometry-aware representations that improve HPE accuracy. We observe that existing HPE models struggle when test images are slightly rotated or flipped, while our method maintains strong performance. Experiments show that our framework matches state-of-the-art models on standard test sets and outperforms them on augmented and full-range poses. Our model handles full-range HPE, accurately predicting head poses across the entire rotation spectrum, including upside-down orientations, and outperforms existing full-yaw range methods.
APA
Wei, T., Hu, H., Liu, H., Wu, X., Fang, Y. & Wu, H.. (2026). CLERF: Contrastive LEaRning for Full-Range Head Pose Estimation. Proceedings of GRaM: the Second Edition of the Workshop on Geometry-grounded Representation Learning and Generative Modeling, in Proceedings of Machine Learning Research 326:519-532 Available from https://proceedings.mlr.press/v326/wei26a.html.

Related Material