EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors

Amin Banayeeanzade, Qingchuan Yang, Deqing Fu, Spencer Hong, Erin Babinsky, Alfy Samuel, Anoop Kumar, Robin Jia, Sai Praneeth Karimireddy
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:6128-6150, 2026.

Abstract

High-quality data is essential for modern machine learning, yet many valuable corpora are sensitive and cannot be freely shared. Synthetic data offers a practical substitute for downstream development, and large language models (LLMs) have emerged as powerful engines for generating it. However, existing private text generation methods are severely inefficient: they are data-intensive, computationally slow, and often require large private corpora or batch sizes to achieve usable quality. We introduce EPSVec, a differentially-private lightweight alternative that steers LLM generation using dataset vectors-directions in activation space that capture the distributional gap between private data and public priors. EPSVec extracts and sanitizes steering vectors just once and then performs standard decoding. This decouples the privacy budget from generation, enabling arbitrarily many synthetic samples without additional privacy cost and yielding strong fidelity even in low-data regimes. Furthermore, we enhance our method by utilizing pretrained (base) models and introducing fixed-shot prompting to boost generation diversity and fidelity. Our experiments demonstrate that EPSVec outperforms existing baselines in distributional alignment and downstream utility, particularly in low-data regimes, while significantly reducing computational overhead.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-banayeeanzade26a, title = {{EPSV}ec: Efficient and Private Synthetic Data Generation via Dataset Vectors}, author = {Banayeeanzade, Amin and Yang, Qingchuan and Fu, Deqing and Hong, Spencer and Babinsky, Erin and Samuel, Alfy and Kumar, Anoop and Jia, Robin and Karimireddy, Sai Praneeth}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {6128--6150}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/banayeeanzade26a/banayeeanzade26a.pdf}, url = {https://proceedings.mlr.press/v306/banayeeanzade26a.html}, abstract = {High-quality data is essential for modern machine learning, yet many valuable corpora are sensitive and cannot be freely shared. Synthetic data offers a practical substitute for downstream development, and large language models (LLMs) have emerged as powerful engines for generating it. However, existing private text generation methods are severely inefficient: they are data-intensive, computationally slow, and often require large private corpora or batch sizes to achieve usable quality. We introduce EPSVec, a differentially-private lightweight alternative that steers LLM generation using dataset vectors-directions in activation space that capture the distributional gap between private data and public priors. EPSVec extracts and sanitizes steering vectors just once and then performs standard decoding. This decouples the privacy budget from generation, enabling arbitrarily many synthetic samples without additional privacy cost and yielding strong fidelity even in low-data regimes. Furthermore, we enhance our method by utilizing pretrained (base) models and introducing fixed-shot prompting to boost generation diversity and fidelity. Our experiments demonstrate that EPSVec outperforms existing baselines in distributional alignment and downstream utility, particularly in low-data regimes, while significantly reducing computational overhead.} }
Endnote
%0 Conference Paper %T EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors %A Amin Banayeeanzade %A Qingchuan Yang %A Deqing Fu %A Spencer Hong %A Erin Babinsky %A Alfy Samuel %A Anoop Kumar %A Robin Jia %A Sai Praneeth Karimireddy %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-banayeeanzade26a %I PMLR %P 6128--6150 %U https://proceedings.mlr.press/v306/banayeeanzade26a.html %V 306 %X High-quality data is essential for modern machine learning, yet many valuable corpora are sensitive and cannot be freely shared. Synthetic data offers a practical substitute for downstream development, and large language models (LLMs) have emerged as powerful engines for generating it. However, existing private text generation methods are severely inefficient: they are data-intensive, computationally slow, and often require large private corpora or batch sizes to achieve usable quality. We introduce EPSVec, a differentially-private lightweight alternative that steers LLM generation using dataset vectors-directions in activation space that capture the distributional gap between private data and public priors. EPSVec extracts and sanitizes steering vectors just once and then performs standard decoding. This decouples the privacy budget from generation, enabling arbitrarily many synthetic samples without additional privacy cost and yielding strong fidelity even in low-data regimes. Furthermore, we enhance our method by utilizing pretrained (base) models and introducing fixed-shot prompting to boost generation diversity and fidelity. Our experiments demonstrate that EPSVec outperforms existing baselines in distributional alignment and downstream utility, particularly in low-data regimes, while significantly reducing computational overhead.
APA
Banayeeanzade, A., Yang, Q., Fu, D., Hong, S., Babinsky, E., Samuel, A., Kumar, A., Jia, R. & Karimireddy, S.P.. (2026). EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:6128-6150 Available from https://proceedings.mlr.press/v306/banayeeanzade26a.html.

Related Material