Composable Parameter-Level Reinforcement Learning for Dual-Chamber Pacemaker Programming

Leqi Jia
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:730-754, 2026.

Abstract

Unnecessary right ventricular pacing (VP) in dual-chamber (DDD) pacemakers is associated with pacing-induced cardiomyopathy, heart failure, and atrial fibrillation. Commercial VP-minimization algorithms such as Medtronic’s Managed Ventricular Pacing (MVP) operate at the beat level and sharply reduce VP burden in patients with intact atrioventricular (AV) conduction (e.g., sick sinus syndrome), but are limited in patients with intermittent AV block, where the algorithm cannot give conduction more time to complete. We propose a hierarchical reinforcement learning framework that operates at the parameter level (AV delay and activity-adaptive target rate) and composes cleanly on top of any certified beat-level controller. Trained via PPO on a clinically parameterized event-level surrogate, the RL policy reduces VP burden from 91.8–95.2% to 5.9–7.8% on Mobitz Type I and Type II when layered on top of standard DDDR. When layered on top of a Casavant-2021-specified MVP controller, it further reduces VP burden from 3.4% to 0.4% on first-degree AV block (driven by AV-delay widening, the same lever as commercial Search-AV) and from 48.6% to 9.2% on Mobitz Type II, while preserving MVP’s near-zero VP (0.2% to 0.1%) on sick sinus syndrome. The MVP implementation is specified against the Casavant–Belk 2021 technical reference and calibrated per subgroup against the IDEAL RVP and COMPARE trials. The parameter-level and beat-level optimizations act on separable timing primitives and combine additively, suggesting a composable deployment path in which a learned parameter-programming layer sits on top of, rather than replaces, certified commercial firmware.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-jia26a, title = {Composable Parameter-Level Reinforcement Learning for Dual-Chamber Pacemaker Programming}, author = {Jia, Leqi}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {730--754}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/jia26a/jia26a.pdf}, url = {https://proceedings.mlr.press/v340/jia26a.html}, abstract = {Unnecessary right ventricular pacing (VP) in dual-chamber (DDD) pacemakers is associated with pacing-induced cardiomyopathy, heart failure, and atrial fibrillation. Commercial VP-minimization algorithms such as Medtronic’s Managed Ventricular Pacing (MVP) operate at the beat level and sharply reduce VP burden in patients with intact atrioventricular (AV) conduction (e.g., sick sinus syndrome), but are limited in patients with intermittent AV block, where the algorithm cannot give conduction more time to complete. We propose a hierarchical reinforcement learning framework that operates at the parameter level (AV delay and activity-adaptive target rate) and composes cleanly on top of any certified beat-level controller. Trained via PPO on a clinically parameterized event-level surrogate, the RL policy reduces VP burden from 91.8–95.2% to 5.9–7.8% on Mobitz Type I and Type II when layered on top of standard DDDR. When layered on top of a Casavant-2021-specified MVP controller, it further reduces VP burden from 3.4% to 0.4% on first-degree AV block (driven by AV-delay widening, the same lever as commercial Search-AV) and from 48.6% to 9.2% on Mobitz Type II, while preserving MVP’s near-zero VP (0.2% to 0.1%) on sick sinus syndrome. The MVP implementation is specified against the Casavant–Belk 2021 technical reference and calibrated per subgroup against the IDEAL RVP and COMPARE trials. The parameter-level and beat-level optimizations act on separable timing primitives and combine additively, suggesting a composable deployment path in which a learned parameter-programming layer sits on top of, rather than replaces, certified commercial firmware.} }
Endnote
%0 Conference Paper %T Composable Parameter-Level Reinforcement Learning for Dual-Chamber Pacemaker Programming %A Leqi Jia %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-jia26a %I PMLR %P 730--754 %U https://proceedings.mlr.press/v340/jia26a.html %V 340 %X Unnecessary right ventricular pacing (VP) in dual-chamber (DDD) pacemakers is associated with pacing-induced cardiomyopathy, heart failure, and atrial fibrillation. Commercial VP-minimization algorithms such as Medtronic’s Managed Ventricular Pacing (MVP) operate at the beat level and sharply reduce VP burden in patients with intact atrioventricular (AV) conduction (e.g., sick sinus syndrome), but are limited in patients with intermittent AV block, where the algorithm cannot give conduction more time to complete. We propose a hierarchical reinforcement learning framework that operates at the parameter level (AV delay and activity-adaptive target rate) and composes cleanly on top of any certified beat-level controller. Trained via PPO on a clinically parameterized event-level surrogate, the RL policy reduces VP burden from 91.8–95.2% to 5.9–7.8% on Mobitz Type I and Type II when layered on top of standard DDDR. When layered on top of a Casavant-2021-specified MVP controller, it further reduces VP burden from 3.4% to 0.4% on first-degree AV block (driven by AV-delay widening, the same lever as commercial Search-AV) and from 48.6% to 9.2% on Mobitz Type II, while preserving MVP’s near-zero VP (0.2% to 0.1%) on sick sinus syndrome. The MVP implementation is specified against the Casavant–Belk 2021 technical reference and calibrated per subgroup against the IDEAL RVP and COMPARE trials. The parameter-level and beat-level optimizations act on separable timing primitives and combine additively, suggesting a composable deployment path in which a learned parameter-programming layer sits on top of, rather than replaces, certified commercial firmware.
APA
Jia, L.. (2026). Composable Parameter-Level Reinforcement Learning for Dual-Chamber Pacemaker Programming. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:730-754 Available from https://proceedings.mlr.press/v340/jia26a.html.

Related Material