Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models

Amir Rezaei Balef, Mykhailo Koshil, Katharina Eggensperger
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:5864-5910, 2026.

Abstract

Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model’s parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-balef26a, title = {Is One Layer Enough? {U}nderstanding Inference Dynamics in Tabular Foundation Models}, author = {Balef, Amir Rezaei and Koshil, Mykhailo and Eggensperger, Katharina}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {5864--5910}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/balef26a/balef26a.pdf}, url = {https://proceedings.mlr.press/v306/balef26a.html}, abstract = {Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model’s parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.} }
Endnote
%0 Conference Paper %T Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models %A Amir Rezaei Balef %A Mykhailo Koshil %A Katharina Eggensperger %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-balef26a %I PMLR %P 5864--5910 %U https://proceedings.mlr.press/v306/balef26a.html %V 306 %X Transformer-based tabular foundation models (TFMs) dominate small to medium tabular predictive benchmark tasks, yet their inference mechanisms remain largely unexplored. We present the first large-scale mechanistic study of layerwise dynamics in 6 state-of-the-art tabular in-context learning models. We explore how predictions emerge across depth, identify distinct stages of inference and reveal latent-space dynamics that differ from those of language models. Our findings indicate substantial depthwise redundancy across multiple models, suggesting iterative refinement with overlapping computations during inference stages. Guided by these insights, we design a proof-of-concept, looped single-layer model that uses only 20% of the original model’s parameters while achieving comparable performance. The code is available at https://github.com/amirbalef/is_one_layer_enough.
APA
Balef, A.R., Koshil, M. & Eggensperger, K.. (2026). Is One Layer Enough? Understanding Inference Dynamics in Tabular Foundation Models. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:5864-5910 Available from https://proceedings.mlr.press/v306/balef26a.html.

Related Material