Collapse-Aware Regularization for Reliable Reasoning Under Distribution Shift

Quynh Vo, Cong-Duy T Nguyen
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:6999-7022, 2026.

Abstract

Reliable reasoning requires models to generalize under distribution shifts, yet in-distribution validation loss often fails to identify checkpoints that remain reliable out of distribution. We study representation collapse in Transformer hidden states as an internal degeneration that may reduce the effective capacity needed for multi-step inference. We characterize two complementary collapse signals: spectral capacity collapse, measured by the effective rank of layerwise representation covariance, and geometric alignment collapse, measured by average cosine alignment. We then propose the Collapse Risk Criterion (CRC), an ID-computable diagnostic score estimated from in-distribution validation representations. CRC is not a formal {OOD} guarantee, but a practical surrogate for {OOD}-free checkpoint selection. Across four reasoning benchmarks and three Transformer backbones, CRC correlates more strongly with {OOD} reasoning error than ID validation loss and standard confidence-based reliability scores. CRC-aware checkpoint selection improves {OOD} performance under depth/difficulty and template/rule shifts with minimal ID degradation. As an extension, collapse-aware training with a CRC-based regularizer improves average {OOD} performance in all evaluated dataset–backbone settings. Our results suggest that monitoring representation collapse is a simple and useful tool for improving reasoning reliability without {OOD} validation data.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-vo26a, title = {Collapse-Aware Regularization for Reliable Reasoning Under Distribution Shift}, author = {Vo, Quynh and Nguyen, Cong-Duy T}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {6999--7022}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/vo26a/vo26a.pdf}, url = {https://proceedings.mlr.press/v337/vo26a.html}, abstract = {Reliable reasoning requires models to generalize under distribution shifts, yet in-distribution validation loss often fails to identify checkpoints that remain reliable out of distribution. We study representation collapse in Transformer hidden states as an internal degeneration that may reduce the effective capacity needed for multi-step inference. We characterize two complementary collapse signals: spectral capacity collapse, measured by the effective rank of layerwise representation covariance, and geometric alignment collapse, measured by average cosine alignment. We then propose the Collapse Risk Criterion (CRC), an ID-computable diagnostic score estimated from in-distribution validation representations. CRC is not a formal {OOD} guarantee, but a practical surrogate for {OOD}-free checkpoint selection. Across four reasoning benchmarks and three Transformer backbones, CRC correlates more strongly with {OOD} reasoning error than ID validation loss and standard confidence-based reliability scores. CRC-aware checkpoint selection improves {OOD} performance under depth/difficulty and template/rule shifts with minimal ID degradation. As an extension, collapse-aware training with a CRC-based regularizer improves average {OOD} performance in all evaluated dataset–backbone settings. Our results suggest that monitoring representation collapse is a simple and useful tool for improving reasoning reliability without {OOD} validation data.} }
Endnote
%0 Conference Paper %T Collapse-Aware Regularization for Reliable Reasoning Under Distribution Shift %A Quynh Vo %A Cong-Duy T Nguyen %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-vo26a %I PMLR %P 6999--7022 %U https://proceedings.mlr.press/v337/vo26a.html %V 337 %X Reliable reasoning requires models to generalize under distribution shifts, yet in-distribution validation loss often fails to identify checkpoints that remain reliable out of distribution. We study representation collapse in Transformer hidden states as an internal degeneration that may reduce the effective capacity needed for multi-step inference. We characterize two complementary collapse signals: spectral capacity collapse, measured by the effective rank of layerwise representation covariance, and geometric alignment collapse, measured by average cosine alignment. We then propose the Collapse Risk Criterion (CRC), an ID-computable diagnostic score estimated from in-distribution validation representations. CRC is not a formal {OOD} guarantee, but a practical surrogate for {OOD}-free checkpoint selection. Across four reasoning benchmarks and three Transformer backbones, CRC correlates more strongly with {OOD} reasoning error than ID validation loss and standard confidence-based reliability scores. CRC-aware checkpoint selection improves {OOD} performance under depth/difficulty and template/rule shifts with minimal ID degradation. As an extension, collapse-aware training with a CRC-based regularizer improves average {OOD} performance in all evaluated dataset–backbone settings. Our results suggest that monitoring representation collapse is a simple and useful tool for improving reasoning reliability without {OOD} validation data.
APA
Vo, Q. & Nguyen, C.T.. (2026). Collapse-Aware Regularization for Reliable Reasoning Under Distribution Shift. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:6999-7022 Available from https://proceedings.mlr.press/v337/vo26a.html.

Related Material