RBCBF: Decoding Time Safety Alignment via Risk Guided Rollback and Barrier Control

Tianxiang Chen, Jingyuan Zhou, Longhao Yan, Kaidi Yang
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:18627-18653, 2026.

Abstract

Existing decoding-time safety interventions are often reactive, relying on local signals to correct unsafe outputs after they emerge. Under adversarial prompts that drive generation into recurring unsafe response, such local signals provide weak guidance for stable repair. As a result, rollback and post-hoc rewriting often trade-off response quality with recurrent violations. To address these limitations, we propose RBCBF, a rollback-based decoding-time framework that jointly selects intervention steps and performs distribution-level corrective control. Our key innovation is a risk-aggregation formulation that views terminal violations as the accumulated build-up of risk along the prefix. By selecting rollback steps from these decisive prefixes, RBCBF moves rollback targeting beyond heuristic cues and turns it into a trajectory-level decision. RBCBF then applies invasive corrective control to the next-token distribution under multiple rule constraints. Across jailbreak-style evaluations, RBCBF outperforms prior rollback methods and decoding-time baselines, reducing harmful responses and substantially lowering violation recurrence.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26ha, title = {{RBCBF}: Decoding Time Safety Alignment via Risk Guided Rollback and Barrier Control}, author = {Chen, Tianxiang and Zhou, Jingyuan and Yan, Longhao and Yang, Kaidi}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {18627--18653}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26ha/chen26ha.pdf}, url = {https://proceedings.mlr.press/v306/chen26ha.html}, abstract = {Existing decoding-time safety interventions are often reactive, relying on local signals to correct unsafe outputs after they emerge. Under adversarial prompts that drive generation into recurring unsafe response, such local signals provide weak guidance for stable repair. As a result, rollback and post-hoc rewriting often trade-off response quality with recurrent violations. To address these limitations, we propose RBCBF, a rollback-based decoding-time framework that jointly selects intervention steps and performs distribution-level corrective control. Our key innovation is a risk-aggregation formulation that views terminal violations as the accumulated build-up of risk along the prefix. By selecting rollback steps from these decisive prefixes, RBCBF moves rollback targeting beyond heuristic cues and turns it into a trajectory-level decision. RBCBF then applies invasive corrective control to the next-token distribution under multiple rule constraints. Across jailbreak-style evaluations, RBCBF outperforms prior rollback methods and decoding-time baselines, reducing harmful responses and substantially lowering violation recurrence.} }
Endnote
%0 Conference Paper %T RBCBF: Decoding Time Safety Alignment via Risk Guided Rollback and Barrier Control %A Tianxiang Chen %A Jingyuan Zhou %A Longhao Yan %A Kaidi Yang %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26ha %I PMLR %P 18627--18653 %U https://proceedings.mlr.press/v306/chen26ha.html %V 306 %X Existing decoding-time safety interventions are often reactive, relying on local signals to correct unsafe outputs after they emerge. Under adversarial prompts that drive generation into recurring unsafe response, such local signals provide weak guidance for stable repair. As a result, rollback and post-hoc rewriting often trade-off response quality with recurrent violations. To address these limitations, we propose RBCBF, a rollback-based decoding-time framework that jointly selects intervention steps and performs distribution-level corrective control. Our key innovation is a risk-aggregation formulation that views terminal violations as the accumulated build-up of risk along the prefix. By selecting rollback steps from these decisive prefixes, RBCBF moves rollback targeting beyond heuristic cues and turns it into a trajectory-level decision. RBCBF then applies invasive corrective control to the next-token distribution under multiple rule constraints. Across jailbreak-style evaluations, RBCBF outperforms prior rollback methods and decoding-time baselines, reducing harmful responses and substantially lowering violation recurrence.
APA
Chen, T., Zhou, J., Yan, L. & Yang, K.. (2026). RBCBF: Decoding Time Safety Alignment via Risk Guided Rollback and Barrier Control. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:18627-18653 Available from https://proceedings.mlr.press/v306/chen26ha.html.

Related Material