Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation

Luoyu Chen, Weiqi Wang, Zhiyi Tian, Chenhan Zhang, Feng Wu, Jianhuan Huang, Ahmed Asiri, Shui Yu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:17165-17189, 2026.

Abstract

Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activations to trigger refusal while preserving benign utility. However, existing steering methods are fundamentally supervised and tied to a static, limited training set, whereas real jailbreaks evolve and are often out-of-distributed from the training set, leading to failures on unseen attacks. In this paper, we tackle this failure by developping a zero-shot defense. Base on unsupervised latent direction discovery, we directly simulate jailbroken activations without any knowledge of jailbreak strategy. To build a defense mechnism upon this, we propose a bi-level adversarial training framework. In the inner step, we simulate diverse jailbroken activations by extrapolating from refusal state harmful-request activations via unsupervised latent direction discovery. In the outer step, we train a potential-induced steering field to push these adversarial jailbroken states into refusal regions while keeping benign unchanged. Across three LLMs and six classical jailbreak families, our method achieves strong defense with attack success rates mostly below 5%, and we analyzed the increasing subspace coverage of our simulated jailbroken activations on real jailbreaks throughout training, which helps explain the increasing robustness of our defense mechnism.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chen26em, title = {Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation}, author = {Chen, Luoyu and Wang, Weiqi and Tian, Zhiyi and Zhang, Chenhan and Wu, Feng and Huang, Jianhuan and Asiri, Ahmed and Yu, Shui}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {17165--17189}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chen26em/chen26em.pdf}, url = {https://proceedings.mlr.press/v306/chen26em.html}, abstract = {Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activations to trigger refusal while preserving benign utility. However, existing steering methods are fundamentally supervised and tied to a static, limited training set, whereas real jailbreaks evolve and are often out-of-distributed from the training set, leading to failures on unseen attacks. In this paper, we tackle this failure by developping a zero-shot defense. Base on unsupervised latent direction discovery, we directly simulate jailbroken activations without any knowledge of jailbreak strategy. To build a defense mechnism upon this, we propose a bi-level adversarial training framework. In the inner step, we simulate diverse jailbroken activations by extrapolating from refusal state harmful-request activations via unsupervised latent direction discovery. In the outer step, we train a potential-induced steering field to push these adversarial jailbroken states into refusal regions while keeping benign unchanged. Across three LLMs and six classical jailbreak families, our method achieves strong defense with attack success rates mostly below 5%, and we analyzed the increasing subspace coverage of our simulated jailbroken activations on real jailbreaks throughout training, which helps explain the increasing robustness of our defense mechnism.} }
Endnote
%0 Conference Paper %T Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation %A Luoyu Chen %A Weiqi Wang %A Zhiyi Tian %A Chenhan Zhang %A Feng Wu %A Jianhuan Huang %A Ahmed Asiri %A Shui Yu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chen26em %I PMLR %P 17165--17189 %U https://proceedings.mlr.press/v306/chen26em.html %V 306 %X Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activations to trigger refusal while preserving benign utility. However, existing steering methods are fundamentally supervised and tied to a static, limited training set, whereas real jailbreaks evolve and are often out-of-distributed from the training set, leading to failures on unseen attacks. In this paper, we tackle this failure by developping a zero-shot defense. Base on unsupervised latent direction discovery, we directly simulate jailbroken activations without any knowledge of jailbreak strategy. To build a defense mechnism upon this, we propose a bi-level adversarial training framework. In the inner step, we simulate diverse jailbroken activations by extrapolating from refusal state harmful-request activations via unsupervised latent direction discovery. In the outer step, we train a potential-induced steering field to push these adversarial jailbroken states into refusal regions while keeping benign unchanged. Across three LLMs and six classical jailbreak families, our method achieves strong defense with attack success rates mostly below 5%, and we analyzed the increasing subspace coverage of our simulated jailbroken activations on real jailbreaks throughout training, which helps explain the increasing robustness of our defense mechnism.
APA
Chen, L., Wang, W., Tian, Z., Zhang, C., Wu, F., Huang, J., Asiri, A. & Yu, S.. (2026). Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:17165-17189 Available from https://proceedings.mlr.press/v306/chen26em.html.

Related Material