Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability

Bum Jun Kim, Shohei Taniguchi, Makoto Kawano, Yusuke Iwasawa, Yutaka Matsuo
Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, PMLR 337:3036-3060, 2026.

Abstract

Training divergence in transformers wastes compute, yet practitioners discover instability only after expensive runs begin. They therefore need an initialization-time signal for identifying risky training configurations. We study Residual {Koopman} Spectral Profiling ({RKSP}), a fixed, single-forward-pass spectral risk score. {RKSP} extracts {Koopman} spectral features by applying whitened dynamic mode decomposition to layer-wise residual snapshots. Our central diagnostic, the near-unit spectral mass, quantifies the fraction of modes concentrated near the unit circle, which captures instability risk. For predicting divergence across extensive configurations, this score achieves an AUROC of 0.995, outperforming the best gradient baseline. When calibrated probabilities are required for a target configuration distribution, a held-out one-dimensional calibrator maps the same scalar score to probabilities. We further make the diagnostic actionable through {Koopman} Spectral Shaping ({KSS}), which reshapes spectra during training. In the challenging high learning rate regime without normalization layers, {KSS} reduces the divergence rate from 66.7% to 12.5% and enables learning rates that are 50% to 150% higher. Controlled experiments establish {RKSP} prediction and {KSS} intervention, real-data language modeling and vision experiments provide transfer evidence, and pretrained language-model profiling, including {GPT-2}, LLaMA-2, and {Qwen3}, demonstrates forward-only diagnostic scalability.

Cite this Paper


BibTeX
@InProceedings{pmlr-v337-kim26d, title = {Residual {Koopman} Spectral Profiling for Predicting and Preventing Transformer Training Instability}, author = {Kim, Bum Jun and Taniguchi, Shohei and Kawano, Makoto and Iwasawa, Yusuke and Matsuo, Yutaka}, booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence}, pages = {3036--3060}, year = {2026}, editor = {Perković, Emilija and Malinsky, Daniel}, volume = {337}, series = {Proceedings of Machine Learning Research}, month = {17--21 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v337/main/assets/kim26d/kim26d.pdf}, url = {https://proceedings.mlr.press/v337/kim26d.html}, abstract = {Training divergence in transformers wastes compute, yet practitioners discover instability only after expensive runs begin. They therefore need an initialization-time signal for identifying risky training configurations. We study Residual {Koopman} Spectral Profiling ({RKSP}), a fixed, single-forward-pass spectral risk score. {RKSP} extracts {Koopman} spectral features by applying whitened dynamic mode decomposition to layer-wise residual snapshots. Our central diagnostic, the near-unit spectral mass, quantifies the fraction of modes concentrated near the unit circle, which captures instability risk. For predicting divergence across extensive configurations, this score achieves an AUROC of 0.995, outperforming the best gradient baseline. When calibrated probabilities are required for a target configuration distribution, a held-out one-dimensional calibrator maps the same scalar score to probabilities. We further make the diagnostic actionable through {Koopman} Spectral Shaping ({KSS}), which reshapes spectra during training. In the challenging high learning rate regime without normalization layers, {KSS} reduces the divergence rate from 66.7% to 12.5% and enables learning rates that are 50% to 150% higher. Controlled experiments establish {RKSP} prediction and {KSS} intervention, real-data language modeling and vision experiments provide transfer evidence, and pretrained language-model profiling, including {GPT-2}, LLaMA-2, and {Qwen3}, demonstrates forward-only diagnostic scalability.} }
Endnote
%0 Conference Paper %T Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability %A Bum Jun Kim %A Shohei Taniguchi %A Makoto Kawano %A Yusuke Iwasawa %A Yutaka Matsuo %B Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2026 %E Emilija Perković %E Daniel Malinsky %F pmlr-v337-kim26d %I PMLR %P 3036--3060 %U https://proceedings.mlr.press/v337/kim26d.html %V 337 %X Training divergence in transformers wastes compute, yet practitioners discover instability only after expensive runs begin. They therefore need an initialization-time signal for identifying risky training configurations. We study Residual {Koopman} Spectral Profiling ({RKSP}), a fixed, single-forward-pass spectral risk score. {RKSP} extracts {Koopman} spectral features by applying whitened dynamic mode decomposition to layer-wise residual snapshots. Our central diagnostic, the near-unit spectral mass, quantifies the fraction of modes concentrated near the unit circle, which captures instability risk. For predicting divergence across extensive configurations, this score achieves an AUROC of 0.995, outperforming the best gradient baseline. When calibrated probabilities are required for a target configuration distribution, a held-out one-dimensional calibrator maps the same scalar score to probabilities. We further make the diagnostic actionable through {Koopman} Spectral Shaping ({KSS}), which reshapes spectra during training. In the challenging high learning rate regime without normalization layers, {KSS} reduces the divergence rate from 66.7% to 12.5% and enables learning rates that are 50% to 150% higher. Controlled experiments establish {RKSP} prediction and {KSS} intervention, real-data language modeling and vision experiments provide transfer evidence, and pretrained language-model profiling, including {GPT-2}, LLaMA-2, and {Qwen3}, demonstrates forward-only diagnostic scalability.
APA
Kim, B.J., Taniguchi, S., Kawano, M., Iwasawa, Y. & Matsuo, Y.. (2026). Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instability. Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research 337:3036-3060 Available from https://proceedings.mlr.press/v337/kim26d.html.

Related Material