[edit]
Composable Parameter-Level Reinforcement Learning for Dual-Chamber Pacemaker Programming
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:730-754, 2026.
Abstract
Unnecessary right ventricular pacing (VP) in dual-chamber (DDD) pacemakers is associated with pacing-induced cardiomyopathy, heart failure, and atrial fibrillation. Commercial VP-minimization algorithms such as Medtronic’s Managed Ventricular Pacing (MVP) operate at the beat level and sharply reduce VP burden in patients with intact atrioventricular (AV) conduction (e.g., sick sinus syndrome), but are limited in patients with intermittent AV block, where the algorithm cannot give conduction more time to complete. We propose a hierarchical reinforcement learning framework that operates at the parameter level (AV delay and activity-adaptive target rate) and composes cleanly on top of any certified beat-level controller. Trained via PPO on a clinically parameterized event-level surrogate, the RL policy reduces VP burden from 91.8–95.2% to 5.9–7.8% on Mobitz Type I and Type II when layered on top of standard DDDR. When layered on top of a Casavant-2021-specified MVP controller, it further reduces VP burden from 3.4% to 0.4% on first-degree AV block (driven by AV-delay widening, the same lever as commercial Search-AV) and from 48.6% to 9.2% on Mobitz Type II, while preserving MVP’s near-zero VP (0.2% to 0.1%) on sick sinus syndrome. The MVP implementation is specified against the Casavant–Belk 2021 technical reference and calibrated per subgroup against the IDEAL RVP and COMPARE trials. The parameter-level and beat-level optimizations act on separable timing primitives and combine additively, suggesting a composable deployment path in which a learned parameter-programming layer sits on top of, rather than replaces, certified commercial firmware.