Efficient LLM Moderation with Multi-Layer Latent Prototypes

Maciej Chrabaszcz, Filip Szatkowski, Bartosz Wójcik, Jan Dubiński, Tomasz Trzcinski, Sebastian Cygert
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:20660-20685, 2026.

Abstract

Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suffer from performance-efficiency trade-offs and are difficult to customize to user-specific requirements. Motivated by this gap, we introduce Multi-Layer Prototype Moderator (MLPM), a lightweight and highly customizable input moderation tool. We propose leveraging prototypes of intermediate representations across multiple layers to improve moderation quality while maintaining high efficiency. By design, our method adds negligible overhead to the generation pipeline and can be seamlessly applied to any model. MLPM achieves state-of-the-art performance on diverse moderation benchmarks and demonstrates strong scalability across model families of various sizes. Moreover, we show that it integrates smoothly into end-to-end moderation pipelines and further improves response safety when combined with output moderation techniques. Overall, our work provides a practical and adaptable solution for safe, robust, and efficient LLM deployment.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chrabaszcz26a, title = {Efficient {LLM} Moderation with Multi-Layer Latent Prototypes}, author = {Chrabaszcz, Maciej and Szatkowski, Filip and W\'{o}jcik, Bartosz and Dubi\'{n}ski, Jan and Trzcinski, Tomasz and Cygert, Sebastian}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {20660--20685}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chrabaszcz26a/chrabaszcz26a.pdf}, url = {https://proceedings.mlr.press/v306/chrabaszcz26a.html}, abstract = {Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suffer from performance-efficiency trade-offs and are difficult to customize to user-specific requirements. Motivated by this gap, we introduce Multi-Layer Prototype Moderator (MLPM), a lightweight and highly customizable input moderation tool. We propose leveraging prototypes of intermediate representations across multiple layers to improve moderation quality while maintaining high efficiency. By design, our method adds negligible overhead to the generation pipeline and can be seamlessly applied to any model. MLPM achieves state-of-the-art performance on diverse moderation benchmarks and demonstrates strong scalability across model families of various sizes. Moreover, we show that it integrates smoothly into end-to-end moderation pipelines and further improves response safety when combined with output moderation techniques. Overall, our work provides a practical and adaptable solution for safe, robust, and efficient LLM deployment.} }
Endnote
%0 Conference Paper %T Efficient LLM Moderation with Multi-Layer Latent Prototypes %A Maciej Chrabaszcz %A Filip Szatkowski %A Bartosz Wójcik %A Jan Dubiński %A Tomasz Trzcinski %A Sebastian Cygert %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chrabaszcz26a %I PMLR %P 20660--20685 %U https://proceedings.mlr.press/v306/chrabaszcz26a.html %V 306 %X Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at deployment time. Existing approaches suffer from performance-efficiency trade-offs and are difficult to customize to user-specific requirements. Motivated by this gap, we introduce Multi-Layer Prototype Moderator (MLPM), a lightweight and highly customizable input moderation tool. We propose leveraging prototypes of intermediate representations across multiple layers to improve moderation quality while maintaining high efficiency. By design, our method adds negligible overhead to the generation pipeline and can be seamlessly applied to any model. MLPM achieves state-of-the-art performance on diverse moderation benchmarks and demonstrates strong scalability across model families of various sizes. Moreover, we show that it integrates smoothly into end-to-end moderation pipelines and further improves response safety when combined with output moderation techniques. Overall, our work provides a practical and adaptable solution for safe, robust, and efficient LLM deployment.
APA
Chrabaszcz, M., Szatkowski, F., Wójcik, B., Dubiński, J., Trzcinski, T. & Cygert, S.. (2026). Efficient LLM Moderation with Multi-Layer Latent Prototypes. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:20660-20685 Available from https://proceedings.mlr.press/v306/chrabaszcz26a.html.

Related Material