When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vilém Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek Suppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Q Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Milind Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:1602-1629, 2026.

Abstract

Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly “saturate”, making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-akhtar26a, title = {When {AI} Benchmarks Plateau: A Systematic Study of Benchmark Saturation}, author = {Akhtar, Mubashara and Reuel, Anka and Soni, Prajna and Ahuja, Sanchit and Ammanamanchi, Pawan Sasanka and Rawal, Ruchit and Zouhar, Vil\'{e}m and Yadav, Srishti and Whitehouse, Chenxi and Ki, Dayeon and Mickel, Jennifer and Choshen, Leshem and Suppa, Marek and Batzner, Jan and Chim, Jenny and Sania, Jeba and Long, Yanan and Rahmani, Hossein A. and Knight, Christina Q and Nan, Yiyang and Raj, Jyoutir and Fan, Yu and Singh, Shubham and Sahoo, Subramanyam and Habba, Eliya and Gohar, Usman and Pawar, Siddhesh Milind and Scholz, Robert and Subramonian, Arjun and Ni, Jingwei and Kochenderfer, Mykel and Koyejo, Sanmi and Sachan, Mrinmaya and Biderman, Stella and Talat, Zeerak and Ghosh, Avijit and Solaiman, Irene}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {1602--1629}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/akhtar26a/akhtar26a.pdf}, url = {https://proceedings.mlr.press/v306/akhtar26a.html}, abstract = {Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly “saturate”, making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches.} }
Endnote
%0 Conference Paper %T When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation %A Mubashara Akhtar %A Anka Reuel %A Prajna Soni %A Sanchit Ahuja %A Pawan Sasanka Ammanamanchi %A Ruchit Rawal %A Vilém Zouhar %A Srishti Yadav %A Chenxi Whitehouse %A Dayeon Ki %A Jennifer Mickel %A Leshem Choshen %A Marek Suppa %A Jan Batzner %A Jenny Chim %A Jeba Sania %A Yanan Long %A Hossein A. Rahmani %A Christina Q Knight %A Yiyang Nan %A Jyoutir Raj %A Yu Fan %A Shubham Singh %A Subramanyam Sahoo %A Eliya Habba %A Usman Gohar %A Siddhesh Milind Pawar %A Robert Scholz %A Arjun Subramonian %A Jingwei Ni %A Mykel Kochenderfer %A Sanmi Koyejo %A Mrinmaya Sachan %A Stella Biderman %A Zeerak Talat %A Avijit Ghosh %A Irene Solaiman %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-akhtar26a %I PMLR %P 1602--1629 %U https://proceedings.mlr.press/v306/akhtar26a.html %V 306 %X Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly “saturate”, making it difficult to differentiate models and diminishing their long-term value. In this study, we define benchmark saturation and analyze it across 60 language model benchmarks using 14 properties that relate to saturation. We find that nearly half of our benchmarks exhibit saturation, with rates increasing with age. Further, we find that resilience to saturation is impacted by expert-curation, not by public test data. Our results suggest that design choices can extend benchmark longevity and inform more durable evaluation approaches.
APA
Akhtar, M., Reuel, A., Soni, P., Ahuja, S., Ammanamanchi, P.S., Rawal, R., Zouhar, V., Yadav, S., Whitehouse, C., Ki, D., Mickel, J., Choshen, L., Suppa, M., Batzner, J., Chim, J., Sania, J., Long, Y., Rahmani, H.A., Knight, C.Q., Nan, Y., Raj, J., Fan, Y., Singh, S., Sahoo, S., Habba, E., Gohar, U., Pawar, S.M., Scholz, R., Subramonian, A., Ni, J., Kochenderfer, M., Koyejo, S., Sachan, M., Biderman, S., Talat, Z., Ghosh, A. & Solaiman, I.. (2026). When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:1602-1629 Available from https://proceedings.mlr.press/v306/akhtar26a.html.

Related Material