Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization

Jinwoo Baek
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:4987-4995, 2026.

Abstract

Low-precision execution can induce substantial forward discrepancies in Transformers even for fixed weights and input, yet these discrepancies are usually monitored only at the output and lack a layer-wise theoretical account. We develop a first-order decomposition of output mismatch into layer-local attention, LayerNorm, and residual-transport terms, and derive from it a practical causal risk estimator and a budgeted controller, Bound-Guided Selective Stabilization (BGSS). Controlled sweeps verify the predicted local sign, monotonicity, and transport structure. On GPT-2, the transport-aware combined predictor is positively correlated with FP32-reference mismatch in all 18 runs and improves over a no-transport ablation in 17/18 runs. Reference-patch attribution shows that the same score preserves useful layer ordering information (mean Spearman 0.362). In budget-matched mitigation, BGSS outperforms random same-budget control in onset events (10.67 vs. 11.67), final mismatch (0.001243 vs. 0.001284), and worst-case mismatch (0.00314 vs. 0.00849), while matching a risk-only same-budget controller on onset suppression and sharply reducing worst-case mismatch (0.00314 vs. 0.00571). These results support a theory-to-algorithm account of Transformer numerical fragility in which finite-precision risk can be analyzed, estimated, localized, and selectively stabilized.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-baek26a, title = { Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization }, author = {Baek, Jinwoo}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {4987--4995}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/baek26a/baek26a.pdf}, url = {https://proceedings.mlr.press/v300/baek26a.html}, abstract = { Low-precision execution can induce substantial forward discrepancies in Transformers even for fixed weights and input, yet these discrepancies are usually monitored only at the output and lack a layer-wise theoretical account. We develop a first-order decomposition of output mismatch into layer-local attention, LayerNorm, and residual-transport terms, and derive from it a practical causal risk estimator and a budgeted controller, Bound-Guided Selective Stabilization (BGSS). Controlled sweeps verify the predicted local sign, monotonicity, and transport structure. On GPT-2, the transport-aware combined predictor is positively correlated with FP32-reference mismatch in all 18 runs and improves over a no-transport ablation in 17/18 runs. Reference-patch attribution shows that the same score preserves useful layer ordering information (mean Spearman 0.362). In budget-matched mitigation, BGSS outperforms random same-budget control in onset events (10.67 vs. 11.67), final mismatch (0.001243 vs. 0.001284), and worst-case mismatch (0.00314 vs. 0.00849), while matching a risk-only same-budget controller on onset suppression and sharply reducing worst-case mismatch (0.00314 vs. 0.00571). These results support a theory-to-algorithm account of Transformer numerical fragility in which finite-precision risk can be analyzed, estimated, localized, and selectively stabilized. } }
Endnote
%0 Conference Paper %T Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization %A Jinwoo Baek %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-baek26a %I PMLR %P 4987--4995 %U https://proceedings.mlr.press/v300/baek26a.html %V 300 %X Low-precision execution can induce substantial forward discrepancies in Transformers even for fixed weights and input, yet these discrepancies are usually monitored only at the output and lack a layer-wise theoretical account. We develop a first-order decomposition of output mismatch into layer-local attention, LayerNorm, and residual-transport terms, and derive from it a practical causal risk estimator and a budgeted controller, Bound-Guided Selective Stabilization (BGSS). Controlled sweeps verify the predicted local sign, monotonicity, and transport structure. On GPT-2, the transport-aware combined predictor is positively correlated with FP32-reference mismatch in all 18 runs and improves over a no-transport ablation in 17/18 runs. Reference-patch attribution shows that the same score preserves useful layer ordering information (mean Spearman 0.362). In budget-matched mitigation, BGSS outperforms random same-budget control in onset events (10.67 vs. 11.67), final mismatch (0.001243 vs. 0.001284), and worst-case mismatch (0.00314 vs. 0.00849), while matching a risk-only same-budget controller on onset suppression and sharply reducing worst-case mismatch (0.00314 vs. 0.00571). These results support a theory-to-algorithm account of Transformer numerical fragility in which finite-precision risk can be analyzed, estimated, localized, and selectively stabilized.
APA
Baek, J.. (2026). Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:4987-4995 Available from https://proceedings.mlr.press/v300/baek26a.html.

Related Material