Adaptive Probe-based Steering for Robust LLM Jailbreaking

Junxi Chen, Junhao Dong, Xiaohua Xie
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:14080-14098, 2026.

Abstract

Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased contrastive prompts and require laborious manual tuning of steering strength, limiting their robustness and effectiveness. In this paper, we leverage the idea of model extraction to guide the learned steering vectors to approximate the ideal one and propose tuning the steering strength adaptively based on contrastive activations’ statistics. Experiments demonstrate that our method notably improves the effectiveness and robustness of probe-based steering, without any extra contrastive prompts or laborious manual tuning. Being an attack paper, this paper focuses on revealing the breakdown of fortified LLMs, raising the average harmfulness score from 6% to 70%. Our code is available at https://github.com/fhdnskfbeuv/adaptiveSteering.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26x, title = {Adaptive Probe-based Steering for Robust {LLM} Jailbreaking}, author = {Chen, Junxi and Dong, Junhao and Xie, Xiaohua}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {14080--14098}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26x/chen26x.pdf}, url = {https://proceedings.mlr.press/v306/chen26x.html}, abstract = {Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased contrastive prompts and require laborious manual tuning of steering strength, limiting their robustness and effectiveness. In this paper, we leverage the idea of model extraction to guide the learned steering vectors to approximate the ideal one and propose tuning the steering strength adaptively based on contrastive activations’ statistics. Experiments demonstrate that our method notably improves the effectiveness and robustness of probe-based steering, without any extra contrastive prompts or laborious manual tuning. Being an attack paper, this paper focuses on revealing the breakdown of fortified LLMs, raising the average harmfulness score from 6% to 70%. Our code is available at https://github.com/fhdnskfbeuv/adaptiveSteering.} }
Endnote
%0 Conference Paper %T Adaptive Probe-based Steering for Robust LLM Jailbreaking %A Junxi Chen %A Junhao Dong %A Xiaohua Xie %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26x %I PMLR %P 14080--14098 %U https://proceedings.mlr.press/v306/chen26x.html %V 306 %X Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased contrastive prompts and require laborious manual tuning of steering strength, limiting their robustness and effectiveness. In this paper, we leverage the idea of model extraction to guide the learned steering vectors to approximate the ideal one and propose tuning the steering strength adaptively based on contrastive activations’ statistics. Experiments demonstrate that our method notably improves the effectiveness and robustness of probe-based steering, without any extra contrastive prompts or laborious manual tuning. Being an attack paper, this paper focuses on revealing the breakdown of fortified LLMs, raising the average harmfulness score from 6% to 70%. Our code is available at https://github.com/fhdnskfbeuv/adaptiveSteering.
APA
Chen, J., Dong, J. & Xie, X.. (2026). Adaptive Probe-based Steering for Robust LLM Jailbreaking. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:14080-14098 Available from https://proceedings.mlr.press/v306/chen26x.html.

Related Material