Detecting Errors in AI-Generated Annotations: When and Why Semantic Neighbors Help

Na Di, Ling Li, Zhe Tang, Hao Cheng, Jinlong Pang, Jiaheng Wei, Zhaowei Zhu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:24622-24653, 2026.

Abstract

Large language models (LLMs) and vision-language models (VLMs) have emerged as efficient annotators for tasks such as generation and classification. While these models offer significant cost and speed advantages over human annotation, a critical challenge remains: existing self-evaluation methods, such as LLM-as-judge, often lack reliable reference-based calibration for error detection. We address this limitation by introducing SAGE (Semantic-Anchored JudGmEnt), a method that leverages semantically similar samples retrieved via $k$-nearest-neighbor as references for annotation verification. We provide a theoretical framework that derives a closed-form expression for the error detection AUROC, which can be decomposed into three factors: intrinsic separability, reference-induced mean shift, and noise reduction through averaging. This decomposition reveals when semantic neighbors help (when references are both semantically matched and correct) and why (by providing reference-based calibration that raises scores for correct annotations and lowers scores for incorrect ones). Experiments on LLM generation, VLM captioning, and classification tasks validate our theoretical framework: SAGE improves error detection when semantic neighbors provide reliable reference-based calibration, and our decomposition offers insights into when direct scoring or alternative strategies may be preferred. Our code is available at https://github.com/dina-1205/SAGE.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-di26b, title = {Detecting Errors in {AI}-Generated Annotations: When and Why Semantic Neighbors Help}, author = {Di, Na and Li, Ling and Tang, Zhe and Cheng, Hao and Pang, Jinlong and Wei, Jiaheng and Zhu, Zhaowei}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {24622--24653}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/di26b/di26b.pdf}, url = {https://proceedings.mlr.press/v306/di26b.html}, abstract = {Large language models (LLMs) and vision-language models (VLMs) have emerged as efficient annotators for tasks such as generation and classification. While these models offer significant cost and speed advantages over human annotation, a critical challenge remains: existing self-evaluation methods, such as LLM-as-judge, often lack reliable reference-based calibration for error detection. We address this limitation by introducing SAGE (Semantic-Anchored JudGmEnt), a method that leverages semantically similar samples retrieved via $k$-nearest-neighbor as references for annotation verification. We provide a theoretical framework that derives a closed-form expression for the error detection AUROC, which can be decomposed into three factors: intrinsic separability, reference-induced mean shift, and noise reduction through averaging. This decomposition reveals when semantic neighbors help (when references are both semantically matched and correct) and why (by providing reference-based calibration that raises scores for correct annotations and lowers scores for incorrect ones). Experiments on LLM generation, VLM captioning, and classification tasks validate our theoretical framework: SAGE improves error detection when semantic neighbors provide reliable reference-based calibration, and our decomposition offers insights into when direct scoring or alternative strategies may be preferred. Our code is available at https://github.com/dina-1205/SAGE.} }
Endnote
%0 Conference Paper %T Detecting Errors in AI-Generated Annotations: When and Why Semantic Neighbors Help %A Na Di %A Ling Li %A Zhe Tang %A Hao Cheng %A Jinlong Pang %A Jiaheng Wei %A Zhaowei Zhu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-di26b %I PMLR %P 24622--24653 %U https://proceedings.mlr.press/v306/di26b.html %V 306 %X Large language models (LLMs) and vision-language models (VLMs) have emerged as efficient annotators for tasks such as generation and classification. While these models offer significant cost and speed advantages over human annotation, a critical challenge remains: existing self-evaluation methods, such as LLM-as-judge, often lack reliable reference-based calibration for error detection. We address this limitation by introducing SAGE (Semantic-Anchored JudGmEnt), a method that leverages semantically similar samples retrieved via $k$-nearest-neighbor as references for annotation verification. We provide a theoretical framework that derives a closed-form expression for the error detection AUROC, which can be decomposed into three factors: intrinsic separability, reference-induced mean shift, and noise reduction through averaging. This decomposition reveals when semantic neighbors help (when references are both semantically matched and correct) and why (by providing reference-based calibration that raises scores for correct annotations and lowers scores for incorrect ones). Experiments on LLM generation, VLM captioning, and classification tasks validate our theoretical framework: SAGE improves error detection when semantic neighbors provide reliable reference-based calibration, and our decomposition offers insights into when direct scoring or alternative strategies may be preferred. Our code is available at https://github.com/dina-1205/SAGE.
APA
Di, N., Li, L., Tang, Z., Cheng, H., Pang, J., Wei, J. & Zhu, Z.. (2026). Detecting Errors in AI-Generated Annotations: When and Why Semantic Neighbors Help. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:24622-24653 Available from https://proceedings.mlr.press/v306/di26b.html.

Related Material