Formally Exploring Time-Series Anomaly Detection Evaluation Metrics

Dennis Wagner, Arjun Nair, Billy Joe Franks, Justus Arweiler, Aparna Muraleedharan, Indra Jungjohann, Fabian Hartung, Andriy Balinskyy, Saurabh Varshneya, Mayank Chetan Ahuja, Nabeel Hussain Syed, Mayank Nagda, Philipp Liznerski, Steffen Reithermann, Maja Rudolph, Sebastian Josef Vollmer, Ralf Schulz, Torsten Katz, Stephan Mandt, Michael Bortz, Heike Leitte, Daniel Neider, Jakob Burger, Fabian Jirasek, Hans Hasse, Sophie Fellenz, Marius Kloft
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:3448-3456, 2026.

Abstract

Detecting anomalies in time series is vital to ensure safety and reliability in many real-world applications. Despite the staggering number of anomaly detection methods, it remains unclear which methods perform best, largely due to flawed evaluation practices. Without rigorous analysis, evaluations yield unintuitive or misleading comparisons. Existing evaluation metrics often focus on specifics and, therefore, fail to capture essential aspects of the anomaly detection task. In this work, we formalize the problem by introducing verifiable properties of evaluation metrics that individually reflect important aspects of anomaly detection in time series. By formalizing requirements and analyzing them systematically, we outline a theoretical framework for evaluating time-series anomaly detection that can support principled evaluations and reliable comparisons. We analyze 37 known metrics and prove that most satisfy only few and none satisfy all properties, explaining many observed inconsistencies in evaluations. To address this gap, we introduce a new flexible evaluation metric LARM that provably satisfies all properties. We illustrate the adaptability of this approach by refining the properties to satisfy stricter requirements and adapting LARM to these advanced properties yielding ALARM.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-wagner26a, title = { Formally Exploring Time-Series Anomaly Detection Evaluation Metrics }, author = {Wagner, Dennis and Nair, Arjun and Franks, Billy Joe and Arweiler, Justus and Muraleedharan, Aparna and Jungjohann, Indra and Hartung, Fabian and Balinskyy, Andriy and Varshneya, Saurabh and Ahuja, Mayank Chetan and Syed, Nabeel Hussain and Nagda, Mayank and Liznerski, Philipp and Reithermann, Steffen and Rudolph, Maja and Vollmer, Sebastian Josef and Schulz, Ralf and Katz, Torsten and Mandt, Stephan and Bortz, Michael and Leitte, Heike and Neider, Daniel and Burger, Jakob and Jirasek, Fabian and Hasse, Hans and Fellenz, Sophie and Kloft, Marius}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {3448--3456}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/wagner26a/wagner26a.pdf}, url = {https://proceedings.mlr.press/v300/wagner26a.html}, abstract = { Detecting anomalies in time series is vital to ensure safety and reliability in many real-world applications. Despite the staggering number of anomaly detection methods, it remains unclear which methods perform best, largely due to flawed evaluation practices. Without rigorous analysis, evaluations yield unintuitive or misleading comparisons. Existing evaluation metrics often focus on specifics and, therefore, fail to capture essential aspects of the anomaly detection task. In this work, we formalize the problem by introducing verifiable properties of evaluation metrics that individually reflect important aspects of anomaly detection in time series. By formalizing requirements and analyzing them systematically, we outline a theoretical framework for evaluating time-series anomaly detection that can support principled evaluations and reliable comparisons. We analyze 37 known metrics and prove that most satisfy only few and none satisfy all properties, explaining many observed inconsistencies in evaluations. To address this gap, we introduce a new flexible evaluation metric LARM that provably satisfies all properties. We illustrate the adaptability of this approach by refining the properties to satisfy stricter requirements and adapting LARM to these advanced properties yielding ALARM. } }
Endnote
%0 Conference Paper %T Formally Exploring Time-Series Anomaly Detection Evaluation Metrics %A Dennis Wagner %A Arjun Nair %A Billy Joe Franks %A Justus Arweiler %A Aparna Muraleedharan %A Indra Jungjohann %A Fabian Hartung %A Andriy Balinskyy %A Saurabh Varshneya %A Mayank Chetan Ahuja %A Nabeel Hussain Syed %A Mayank Nagda %A Philipp Liznerski %A Steffen Reithermann %A Maja Rudolph %A Sebastian Josef Vollmer %A Ralf Schulz %A Torsten Katz %A Stephan Mandt %A Michael Bortz %A Heike Leitte %A Daniel Neider %A Jakob Burger %A Fabian Jirasek %A Hans Hasse %A Sophie Fellenz %A Marius Kloft %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-wagner26a %I PMLR %P 3448--3456 %U https://proceedings.mlr.press/v300/wagner26a.html %V 300 %X Detecting anomalies in time series is vital to ensure safety and reliability in many real-world applications. Despite the staggering number of anomaly detection methods, it remains unclear which methods perform best, largely due to flawed evaluation practices. Without rigorous analysis, evaluations yield unintuitive or misleading comparisons. Existing evaluation metrics often focus on specifics and, therefore, fail to capture essential aspects of the anomaly detection task. In this work, we formalize the problem by introducing verifiable properties of evaluation metrics that individually reflect important aspects of anomaly detection in time series. By formalizing requirements and analyzing them systematically, we outline a theoretical framework for evaluating time-series anomaly detection that can support principled evaluations and reliable comparisons. We analyze 37 known metrics and prove that most satisfy only few and none satisfy all properties, explaining many observed inconsistencies in evaluations. To address this gap, we introduce a new flexible evaluation metric LARM that provably satisfies all properties. We illustrate the adaptability of this approach by refining the properties to satisfy stricter requirements and adapting LARM to these advanced properties yielding ALARM.
APA
Wagner, D., Nair, A., Franks, B.J., Arweiler, J., Muraleedharan, A., Jungjohann, I., Hartung, F., Balinskyy, A., Varshneya, S., Ahuja, M.C., Syed, N.H., Nagda, M., Liznerski, P., Reithermann, S., Rudolph, M., Vollmer, S.J., Schulz, R., Katz, T., Mandt, S., Bortz, M., Leitte, H., Neider, D., Burger, J., Jirasek, F., Hasse, H., Fellenz, S. & Kloft, M.. (2026). Formally Exploring Time-Series Anomaly Detection Evaluation Metrics . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:3448-3456 Available from https://proceedings.mlr.press/v300/wagner26a.html.

Related Material