CTRLS: Chain-of-Thought Reasoning via Latent State Transition

Junda Wu, Yuxin Xiong, Xintong Li, Sheldon Yu, Zhengmian Hu, Tong Yu, Rui Wang, Xiang Chen, Jingbo Shang, Julian McAuley
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:2728-2736, 2026.

Abstract

Chain-of-thought (CoT) reasoning enables large language models (LLMs) to break down complex problems into explainable intermediate steps, significantly enhancing model transparency and performance in reasoning tasks. However, conventional CoT methods rely on heuristic sampling without structured modelling of reasoning transitions, constraining their ability to explore and discover diverse and effective reasoning trajectories. In this work, we introduce CTRLS, a framework that formulates CoT reasoning as a Markov decision process (MDP) with latent state transitions, enabling explainable and state-aware exploration via distributional reinforcement learning. By modelling reasoning actions as explicit probability distributions in latent space, our approach explicitly models epistemic uncertainty, facilitating robust exploration of the reasoning space. Enabled by our formulation, we propose an on-policy reinforcement learning scheme to iteratively refine latent transitions without fine-tuning of the underlying LLM. Theoretical analyses provide evidence lower bounds (ELBO), theoretically grounding our transition-aware modelling of latent reasoning dynamics.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-wu26a, title = { CTRLS: Chain-of-Thought Reasoning via Latent State Transition }, author = {Wu, Junda and Xiong, Yuxin and Li, Xintong and Yu, Sheldon and Hu, Zhengmian and Yu, Tong and Wang, Rui and Chen, Xiang and Shang, Jingbo and McAuley, Julian}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {2728--2736}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/wu26a/wu26a.pdf}, url = {https://proceedings.mlr.press/v300/wu26a.html}, abstract = { Chain-of-thought (CoT) reasoning enables large language models (LLMs) to break down complex problems into explainable intermediate steps, significantly enhancing model transparency and performance in reasoning tasks. However, conventional CoT methods rely on heuristic sampling without structured modelling of reasoning transitions, constraining their ability to explore and discover diverse and effective reasoning trajectories. In this work, we introduce CTRLS, a framework that formulates CoT reasoning as a Markov decision process (MDP) with latent state transitions, enabling explainable and state-aware exploration via distributional reinforcement learning. By modelling reasoning actions as explicit probability distributions in latent space, our approach explicitly models epistemic uncertainty, facilitating robust exploration of the reasoning space. Enabled by our formulation, we propose an on-policy reinforcement learning scheme to iteratively refine latent transitions without fine-tuning of the underlying LLM. Theoretical analyses provide evidence lower bounds (ELBO), theoretically grounding our transition-aware modelling of latent reasoning dynamics. } }
Endnote
%0 Conference Paper %T CTRLS: Chain-of-Thought Reasoning via Latent State Transition %A Junda Wu %A Yuxin Xiong %A Xintong Li %A Sheldon Yu %A Zhengmian Hu %A Tong Yu %A Rui Wang %A Xiang Chen %A Jingbo Shang %A Julian McAuley %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-wu26a %I PMLR %P 2728--2736 %U https://proceedings.mlr.press/v300/wu26a.html %V 300 %X Chain-of-thought (CoT) reasoning enables large language models (LLMs) to break down complex problems into explainable intermediate steps, significantly enhancing model transparency and performance in reasoning tasks. However, conventional CoT methods rely on heuristic sampling without structured modelling of reasoning transitions, constraining their ability to explore and discover diverse and effective reasoning trajectories. In this work, we introduce CTRLS, a framework that formulates CoT reasoning as a Markov decision process (MDP) with latent state transitions, enabling explainable and state-aware exploration via distributional reinforcement learning. By modelling reasoning actions as explicit probability distributions in latent space, our approach explicitly models epistemic uncertainty, facilitating robust exploration of the reasoning space. Enabled by our formulation, we propose an on-policy reinforcement learning scheme to iteratively refine latent transitions without fine-tuning of the underlying LLM. Theoretical analyses provide evidence lower bounds (ELBO), theoretically grounding our transition-aware modelling of latent reasoning dynamics.
APA
Wu, J., Xiong, Y., Li, X., Yu, S., Hu, Z., Yu, T., Wang, R., Chen, X., Shang, J. & McAuley, J.. (2026). CTRLS: Chain-of-Thought Reasoning via Latent State Transition . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:2728-2736 Available from https://proceedings.mlr.press/v300/wu26a.html.

Related Material