Separating Representation from Reconstruction Enables Scalable Text Encoders

Megi Dervishi, Mathurin Videau, Yann Lecun
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:24500-24514, 2026.

Abstract

While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the representations of BERT encoders become increasingly unexploitable by frozen probes, despite improved perplexity. The misalignment originates in BERT’s flat design, which couples representation learning to the token reconstruction loss. We propose CrossBERT, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction. This design further enables high masking ratios ($\geq 50$%) and gradient collection over all tokens via a Complementary Masking Strategy, respectively increasing throughput by $1.5$ to $2\times$ and sample efficiency by $2\times$. Overall, CrossBERT demonstrates monotonic scaling and superior performance on MTEB(eng, v2) and frozen GLUE benchmarks.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-dervishi26a, title = {Separating Representation from Reconstruction Enables Scalable Text Encoders}, author = {Dervishi, Megi and Videau, Mathurin and Lecun, Yann}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {24500--24514}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/dervishi26a/dervishi26a.pdf}, url = {https://proceedings.mlr.press/v306/dervishi26a.html}, abstract = {While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the representations of BERT encoders become increasingly unexploitable by frozen probes, despite improved perplexity. The misalignment originates in BERT’s flat design, which couples representation learning to the token reconstruction loss. We propose CrossBERT, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction. This design further enables high masking ratios ($\geq 50$%) and gradient collection over all tokens via a Complementary Masking Strategy, respectively increasing throughput by $1.5$ to $2\times$ and sample efficiency by $2\times$. Overall, CrossBERT demonstrates monotonic scaling and superior performance on MTEB(eng, v2) and frozen GLUE benchmarks.} }
Endnote
%0 Conference Paper %T Separating Representation from Reconstruction Enables Scalable Text Encoders %A Megi Dervishi %A Mathurin Videau %A Yann Lecun %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-dervishi26a %I PMLR %P 24500--24514 %U https://proceedings.mlr.press/v306/dervishi26a.html %V 306 %X While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evaluation via probing. Under this lens, the representations of BERT encoders become increasingly unexploitable by frozen probes, despite improved perplexity. The misalignment originates in BERT’s flat design, which couples representation learning to the token reconstruction loss. We propose CrossBERT, a two-part architecture that separates the learning of high-quality encoded representations from the rigid grounding of token reconstruction. This design further enables high masking ratios ($\geq 50$%) and gradient collection over all tokens via a Complementary Masking Strategy, respectively increasing throughput by $1.5$ to $2\times$ and sample efficiency by $2\times$. Overall, CrossBERT demonstrates monotonic scaling and superior performance on MTEB(eng, v2) and frozen GLUE benchmarks.
APA
Dervishi, M., Videau, M. & Lecun, Y.. (2026). Separating Representation from Reconstruction Enables Scalable Text Encoders. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:24500-24514 Available from https://proceedings.mlr.press/v306/dervishi26a.html.

Related Material