Dataset Distillation with Convexified Implicit Gradients

Noel Loo; Ramin Hasani; Mathias Lechner; Daniela Rus

Dataset Distillation with Convexified Implicit Gradients

Noel Loo, Ramin Hasani, Mathias Lechner, Daniela Rus

Proceedings of the 40th International Conference on Machine Learning, PMLR 202:22649-22674, 2023.

Abstract

We propose a new dataset distillation algorithm using reparameterization and convexification of implicit gradients (RCIG), that substantially improves the state-of-the-art. To this end, we first formulate dataset distillation as a bi-level optimization problem. Then, we show how implicit gradients can be effectively used to compute meta-gradient updates. We further equip the algorithm with a convexified approximation that corresponds to learning on top of a frozen finite-width neural tangent kernel. Finally, we improve bias in implicit gradients by parameterizing the neural network to enable analytical computation of final-layer parameters given the body parameters. RCIG establishes the new state-of-the-art on a diverse series of dataset distillation tasks. Notably, with one image per class, on resized ImageNet, RCIG sees on average a 108% improvement over the previous state-of-the-art distillation algorithm. Similarly, we observed a 66% gain over SOTA on Tiny-ImageNet and 37% on CIFAR-100.

Cite this Paper

BibTeX

@InProceedings{pmlr-v202-loo23a,
  title = 	 {Dataset Distillation with Convexified Implicit Gradients},
  author =       {Loo, Noel and Hasani, Ramin and Lechner, Mathias and Rus, Daniela},
  booktitle = 	 {Proceedings of the 40th International Conference on Machine Learning},
  pages = 	 {22649--22674},
  year = 	 {2023},
  editor = 	 {Krause, Andreas and Brunskill, Emma and Cho, Kyunghyun and Engelhardt, Barbara and Sabato, Sivan and Scarlett, Jonathan},
  volume = 	 {202},
  series = 	 {Proceedings of Machine Learning Research},
  month = 	 {23--29 Jul},
  publisher =    {PMLR},
  pdf = 	 {https://proceedings.mlr.press/v202/loo23a/loo23a.pdf},
  url = 	 {https://proceedings.mlr.press/v202/loo23a.html},
  abstract = 	 {We propose a new dataset distillation algorithm using reparameterization and convexification of implicit gradients (RCIG), that substantially improves the state-of-the-art. To this end, we first formulate dataset distillation as a bi-level optimization problem. Then, we show how implicit gradients can be effectively used to compute meta-gradient updates. We further equip the algorithm with a convexified approximation that corresponds to learning on top of a frozen finite-width neural tangent kernel. Finally, we improve bias in implicit gradients by parameterizing the neural network to enable analytical computation of final-layer parameters given the body parameters. RCIG establishes the new state-of-the-art on a diverse series of dataset distillation tasks. Notably, with one image per class, on resized ImageNet, RCIG sees on average a 108% improvement over the previous state-of-the-art distillation algorithm. Similarly, we observed a 66% gain over SOTA on Tiny-ImageNet and 37% on CIFAR-100.}
}

Endnote

%0 Conference Paper
%T Dataset Distillation with Convexified Implicit Gradients
%A Noel Loo
%A Ramin Hasani
%A Mathias Lechner
%A Daniela Rus
%B Proceedings of the 40th International Conference on Machine Learning
%C Proceedings of Machine Learning Research
%D 2023
%E Andreas Krause
%E Emma Brunskill
%E Kyunghyun Cho
%E Barbara Engelhardt
%E Sivan Sabato
%E Jonathan Scarlett	
%F pmlr-v202-loo23a
%I PMLR
%P 22649--22674
%U https://proceedings.mlr.press/v202/loo23a.html
%V 202
%X We propose a new dataset distillation algorithm using reparameterization and convexification of implicit gradients (RCIG), that substantially improves the state-of-the-art. To this end, we first formulate dataset distillation as a bi-level optimization problem. Then, we show how implicit gradients can be effectively used to compute meta-gradient updates. We further equip the algorithm with a convexified approximation that corresponds to learning on top of a frozen finite-width neural tangent kernel. Finally, we improve bias in implicit gradients by parameterizing the neural network to enable analytical computation of final-layer parameters given the body parameters. RCIG establishes the new state-of-the-art on a diverse series of dataset distillation tasks. Notably, with one image per class, on resized ImageNet, RCIG sees on average a 108% improvement over the previous state-of-the-art distillation algorithm. Similarly, we observed a 66% gain over SOTA on Tiny-ImageNet and 37% on CIFAR-100.

APA

Loo, N., Hasani, R., Lechner, M. & Rus, D.. (2023). Dataset Distillation with Convexified Implicit Gradients. Proceedings of the 40th International Conference on Machine Learning, in Proceedings of Machine Learning Research 202:22649-22674 Available from https://proceedings.mlr.press/v202/loo23a.html.

Dataset Distillation with Convexified Implicit Gradients

Abstract

Cite this Paper

Related Material