Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces

Anjiang Wei; Allen Nie; Thiago S. F. X. Teixeira; Rohan Yadav; Wonchan Lee; Ke Wang; Alex Aiken

Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces

Anjiang Wei, Allen Nie, Thiago S. F. X. Teixeira, Rohan Yadav, Wonchan Lee, Ke Wang, Alex Aiken

Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:66155-66177, 2025.

Abstract

Modern scientific discovery increasingly relies on high-performance computing for complex modeling and simulation. A key challenge in improving parallel program performance is efficiently mapping tasks to processors and data to memory, a process dictated by intricate, low-level system code known as mappers. Developing high-performance mappers demands days of manual tuning, posing a significant barrier for domain scientists without systems expertise. We introduce a framework that automates mapper development with generative optimization, leveraging richer feedback beyond scalar performance metrics. Our approach features the Agent-System Interface, which includes a Domain-Specific Language (DSL) to abstract away the low-level complexity of system code and define a structured search space, as well as AutoGuide, a mechanism that interprets raw execution output into actionable feedback. Unlike traditional reinforcement learning methods such as OpenTuner, which rely solely on scalar feedback, our method finds superior mappers in far fewer iterations. With just 10 iterations, it outperforms OpenTuner even after 1000 iterations, achieving $3.8\times$ faster performance. Our approach finds mappers that surpass expert-written mappers by up to $1.34\times$ speedup across nine benchmarks while reducing tuning time from days to minutes.

Cite this Paper

BibTeX

@InProceedings{pmlr-v267-wei25j,
  title = 	 {Improving Parallel Program Performance with {LLM} Optimizers via Agent-System Interfaces},
  author =       {Wei, Anjiang and Nie, Allen and Teixeira, Thiago S. F. X. and Yadav, Rohan and Lee, Wonchan and Wang, Ke and Aiken, Alex},
  booktitle = 	 {Proceedings of the 42nd International Conference on Machine Learning},
  pages = 	 {66155--66177},
  year = 	 {2025},
  editor = 	 {Singh, Aarti and Fazel, Maryam and Hsu, Daniel and Lacoste-Julien, Simon and Berkenkamp, Felix and Maharaj, Tegan and Wagstaff, Kiri and Zhu, Jerry},
  volume = 	 {267},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {13--19 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://raw.githubusercontent.com/mlresearch/v267/main/assets/wei25j/wei25j.pdf},
  url = 	 {https://proceedings.mlr.press/v267/wei25j.html},
  abstract = 	 {Modern scientific discovery increasingly relies on high-performance computing for complex modeling and simulation. A key challenge in improving parallel program performance is efficiently mapping tasks to processors and data to memory, a process dictated by intricate, low-level system code known as mappers. Developing high-performance mappers demands days of manual tuning, posing a significant barrier for domain scientists without systems expertise. We introduce a framework that automates mapper development with generative optimization, leveraging richer feedback beyond scalar performance metrics. Our approach features the Agent-System Interface, which includes a Domain-Specific Language (DSL) to abstract away the low-level complexity of system code and define a structured search space, as well as AutoGuide, a mechanism that interprets raw execution output into actionable feedback. Unlike traditional reinforcement learning methods such as OpenTuner, which rely solely on scalar feedback, our method finds superior mappers in far fewer iterations. With just 10 iterations, it outperforms OpenTuner even after 1000 iterations, achieving $3.8\times$ faster performance. Our approach finds mappers that surpass expert-written mappers by up to $1.34\times$ speedup across nine benchmarks while reducing tuning time from days to minutes.}
}

Endnote

%0 Conference Paper
%T Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces
%A Anjiang Wei
%A Allen Nie
%A Thiago S. F. X. Teixeira
%A Rohan Yadav
%A Wonchan Lee
%A Ke Wang
%A Alex Aiken
%B Proceedings of the 42nd International Conference on Machine Learning
%C Proceedings of Machine Learning Research
%D 2025
%E Aarti Singh
%E Maryam Fazel
%E Daniel Hsu
%E Simon Lacoste-Julien
%E Felix Berkenkamp
%E Tegan Maharaj
%E Kiri Wagstaff
%E Jerry Zhu	
%F pmlr-v267-wei25j
%I PMLR
%P 66155--66177
%U https://proceedings.mlr.press/v267/wei25j.html
%V 267
%X Modern scientific discovery increasingly relies on high-performance computing for complex modeling and simulation. A key challenge in improving parallel program performance is efficiently mapping tasks to processors and data to memory, a process dictated by intricate, low-level system code known as mappers. Developing high-performance mappers demands days of manual tuning, posing a significant barrier for domain scientists without systems expertise. We introduce a framework that automates mapper development with generative optimization, leveraging richer feedback beyond scalar performance metrics. Our approach features the Agent-System Interface, which includes a Domain-Specific Language (DSL) to abstract away the low-level complexity of system code and define a structured search space, as well as AutoGuide, a mechanism that interprets raw execution output into actionable feedback. Unlike traditional reinforcement learning methods such as OpenTuner, which rely solely on scalar feedback, our method finds superior mappers in far fewer iterations. With just 10 iterations, it outperforms OpenTuner even after 1000 iterations, achieving $3.8\times$ faster performance. Our approach finds mappers that surpass expert-written mappers by up to $1.34\times$ speedup across nine benchmarks while reducing tuning time from days to minutes.

APA

Wei, A., Nie, A., Teixeira, T.S.F.X., Yadav, R., Lee, W., Wang, K. & Aiken, A.. (2025). Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces. Proceedings of the 42nd International Conference on Machine Learning, in Proceedings of Machine Learning Research 267:66155-66177 Available from https://proceedings.mlr.press/v267/wei25j.html.

Improving Parallel Program Performance with LLM Optimizers via Agent-System Interfaces

Abstract

Cite this Paper

Related Material