Learning Rewrite-Invariant Reasoning with Targeted Alternation Training

Mousa Arraf, Ido Guy, Kira Radinsky
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:3796-3818, 2026.

Abstract

Large language models (LLMs) often fail in systematic, model-specific ways under meaning-preserving question rewrites (paraphrases, format changes, benign distractors). In this work, we address this instability by identifying where the model’s reasoning process diverges across semantically-equivalent inputs. For each target LLM, we sample multiple solution traces under rewrites and aggregate them into a graph of recurring intermediate steps, which pinpoints where incorrect traces diverge from correct ones. We then generate a small set of semantics-preserving examples that mirror the rewrite patterns most responsible for these divergences, and use them to steer the model (targeted alternation training), either via fine-tuning or via in-context learning. Across MMLU-Pro, Big-MATH, and DROP, this yields consistent gains and cross-dataset generalization. On Humanity’s Last Exam, using 200 in-context examples, it improves GPT-5.2 (xhigh) from 35.4% to 38.1%, demonstrating that targeted alternation training can materially improve a frontier, API-accessible closed model under realistic access constraints.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-arraf26a, title = {Learning Rewrite-Invariant Reasoning with Targeted Alternation Training}, author = {Arraf, Mousa and Guy, Ido and Radinsky, Kira}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {3796--3818}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/arraf26a/arraf26a.pdf}, url = {https://proceedings.mlr.press/v306/arraf26a.html}, abstract = {Large language models (LLMs) often fail in systematic, model-specific ways under meaning-preserving question rewrites (paraphrases, format changes, benign distractors). In this work, we address this instability by identifying where the model’s reasoning process diverges across semantically-equivalent inputs. For each target LLM, we sample multiple solution traces under rewrites and aggregate them into a graph of recurring intermediate steps, which pinpoints where incorrect traces diverge from correct ones. We then generate a small set of semantics-preserving examples that mirror the rewrite patterns most responsible for these divergences, and use them to steer the model (targeted alternation training), either via fine-tuning or via in-context learning. Across MMLU-Pro, Big-MATH, and DROP, this yields consistent gains and cross-dataset generalization. On Humanity’s Last Exam, using 200 in-context examples, it improves GPT-5.2 (xhigh) from 35.4% to 38.1%, demonstrating that targeted alternation training can materially improve a frontier, API-accessible closed model under realistic access constraints.} }
Endnote
%0 Conference Paper %T Learning Rewrite-Invariant Reasoning with Targeted Alternation Training %A Mousa Arraf %A Ido Guy %A Kira Radinsky %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-arraf26a %I PMLR %P 3796--3818 %U https://proceedings.mlr.press/v306/arraf26a.html %V 306 %X Large language models (LLMs) often fail in systematic, model-specific ways under meaning-preserving question rewrites (paraphrases, format changes, benign distractors). In this work, we address this instability by identifying where the model’s reasoning process diverges across semantically-equivalent inputs. For each target LLM, we sample multiple solution traces under rewrites and aggregate them into a graph of recurring intermediate steps, which pinpoints where incorrect traces diverge from correct ones. We then generate a small set of semantics-preserving examples that mirror the rewrite patterns most responsible for these divergences, and use them to steer the model (targeted alternation training), either via fine-tuning or via in-context learning. Across MMLU-Pro, Big-MATH, and DROP, this yields consistent gains and cross-dataset generalization. On Humanity’s Last Exam, using 200 in-context examples, it improves GPT-5.2 (xhigh) from 35.4% to 38.1%, demonstrating that targeted alternation training can materially improve a frontier, API-accessible closed model under realistic access constraints.
APA
Arraf, M., Guy, I. & Radinsky, K.. (2026). Learning Rewrite-Invariant Reasoning with Targeted Alternation Training. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:3796-3818 Available from https://proceedings.mlr.press/v306/arraf26a.html.

Related Material