Don’t Walk the Line: Boundary Guidance for Filtered Generation

Sarah Ball, Andreas Haupt
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:6024-6040, 2026.

Abstract

Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier’s decision boundary, increasing both false positives and false negatives. We propose Boundary Guidance, a reinforcement learning fine-tuning method that explicitly steers generation away from the classifier’s margin. On a benchmark of jailbreak, ambiguous, and long-context prompts, Boundary Guidance improves the safety while maintaining or improving the utility of outputs, as judged by LLM-as-a-Judge evaluations. Comprehensive ablations across model scales and reward designs demonstrate the robustness of our approach.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-ball26b, title = {Don’t Walk the Line: Boundary Guidance for Filtered Generation}, author = {Ball, Sarah and Haupt, Andreas}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {6024--6040}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ball26b/ball26b.pdf}, url = {https://proceedings.mlr.press/v306/ball26b.html}, abstract = {Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier’s decision boundary, increasing both false positives and false negatives. We propose Boundary Guidance, a reinforcement learning fine-tuning method that explicitly steers generation away from the classifier’s margin. On a benchmark of jailbreak, ambiguous, and long-context prompts, Boundary Guidance improves the safety while maintaining or improving the utility of outputs, as judged by LLM-as-a-Judge evaluations. Comprehensive ablations across model scales and reward designs demonstrate the robustness of our approach.} }
Endnote
%0 Conference Paper %T Don’t Walk the Line: Boundary Guidance for Filtered Generation %A Sarah Ball %A Andreas Haupt %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-ball26b %I PMLR %P 6024--6040 %U https://proceedings.mlr.press/v306/ball26b.html %V 306 %X Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tune the generator to reduce the probability of being filtered, but this can be suboptimal: it often pushes the model toward producing samples near the classifier’s decision boundary, increasing both false positives and false negatives. We propose Boundary Guidance, a reinforcement learning fine-tuning method that explicitly steers generation away from the classifier’s margin. On a benchmark of jailbreak, ambiguous, and long-context prompts, Boundary Guidance improves the safety while maintaining or improving the utility of outputs, as judged by LLM-as-a-Judge evaluations. Comprehensive ablations across model scales and reward designs demonstrate the robustness of our approach.
APA
Ball, S. & Haupt, A.. (2026). Don’t Walk the Line: Boundary Guidance for Filtered Generation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:6024-6040 Available from https://proceedings.mlr.press/v306/ball26b.html.

Related Material