Annealing Paths for the Evaluation of Topic Models

James Foulds, Padhraic Smyth
Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, PMLR R12:17-26, 2014.

Abstract

Statistical topic models such as latent Dirich- let allocation have become enormously popu- lar in the past decade, with dozens of learning algorithms and extensions being proposed each year. As these models and algorithms continue to be developed, it becomes increasingly impor- tant to evaluate them relative to previous tech- niques. However, evaluating the predictive per- formance of a topic model is a computationally difficult task. Annealed importance sampling (AIS), a Monte Carlo technique which operates by annealing between two distributions, has pre- viously been successfully used for topic model evaluation (Wallach et al., 2009b). This tech- nique estimates the likelihood of a held-out doc- ument by simulating an annealing process from the prior to the posterior for the latent topic as- signments, and using this simulation as an im- portance sampling proposal distribution. In this paper we introduce new AIS annealing paths which instead anneal from one topic model to another, thereby estimating the relative perfor- mance of the models. This strategy can exhibit much lower empirical variance than previous ap- proaches, facilitating reliable per-documentcom- parisons of topic models. We then show how to use these paths to evaluate the predictive perfor- mance of topic model learning algorithms by effi- ciently estimating the likelihood at each iteration of the training procedure. The proposed method achieves better held-out likelihood estimates for this task than previous algorithms with, in some cases, an order of magnitude less computation.

Cite this Paper


BibTeX
@InProceedings{pmlr-vR12-foulds14a, title = {Annealing Paths for the Evaluation of Topic Models}, author = {Foulds, James and Smyth, Padhraic}, booktitle = {Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence}, pages = {17--26}, year = {2014}, editor = {Zhang, Nevin L. and Tian, Jin}, volume = {R12}, series = {Proceedings of Machine Learning Research}, month = {23--27 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/r12/main/assets/foulds14a/foulds14a.pdf}, url = {https://proceedings.mlr.press/r12/foulds14a.html}, abstract = {Statistical topic models such as latent Dirich- let allocation have become enormously popu- lar in the past decade, with dozens of learning algorithms and extensions being proposed each year. As these models and algorithms continue to be developed, it becomes increasingly impor- tant to evaluate them relative to previous tech- niques. However, evaluating the predictive per- formance of a topic model is a computationally difficult task. Annealed importance sampling (AIS), a Monte Carlo technique which operates by annealing between two distributions, has pre- viously been successfully used for topic model evaluation (Wallach et al., 2009b). This tech- nique estimates the likelihood of a held-out doc- ument by simulating an annealing process from the prior to the posterior for the latent topic as- signments, and using this simulation as an im- portance sampling proposal distribution. In this paper we introduce new AIS annealing paths which instead anneal from one topic model to another, thereby estimating the relative perfor- mance of the models. This strategy can exhibit much lower empirical variance than previous ap- proaches, facilitating reliable per-documentcom- parisons of topic models. We then show how to use these paths to evaluate the predictive perfor- mance of topic model learning algorithms by effi- ciently estimating the likelihood at each iteration of the training procedure. The proposed method achieves better held-out likelihood estimates for this task than previous algorithms with, in some cases, an order of magnitude less computation.}, note = {Reissued by PMLR on 04 October 2026.} }
Endnote
%0 Conference Paper %T Annealing Paths for the Evaluation of Topic Models %A James Foulds %A Padhraic Smyth %B Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence %C Proceedings of Machine Learning Research %D 2014 %E Nevin L. Zhang %E Jin Tian %F pmlr-vR12-foulds14a %I PMLR %P 17--26 %U https://proceedings.mlr.press/r12/foulds14a.html %V R12 %X Statistical topic models such as latent Dirich- let allocation have become enormously popu- lar in the past decade, with dozens of learning algorithms and extensions being proposed each year. As these models and algorithms continue to be developed, it becomes increasingly impor- tant to evaluate them relative to previous tech- niques. However, evaluating the predictive per- formance of a topic model is a computationally difficult task. Annealed importance sampling (AIS), a Monte Carlo technique which operates by annealing between two distributions, has pre- viously been successfully used for topic model evaluation (Wallach et al., 2009b). This tech- nique estimates the likelihood of a held-out doc- ument by simulating an annealing process from the prior to the posterior for the latent topic as- signments, and using this simulation as an im- portance sampling proposal distribution. In this paper we introduce new AIS annealing paths which instead anneal from one topic model to another, thereby estimating the relative perfor- mance of the models. This strategy can exhibit much lower empirical variance than previous ap- proaches, facilitating reliable per-documentcom- parisons of topic models. We then show how to use these paths to evaluate the predictive perfor- mance of topic model learning algorithms by effi- ciently estimating the likelihood at each iteration of the training procedure. The proposed method achieves better held-out likelihood estimates for this task than previous algorithms with, in some cases, an order of magnitude less computation. %Z Reissued by PMLR on 04 October 2026.
APA
Foulds, J. & Smyth, P.. (2014). Annealing Paths for the Evaluation of Topic Models. Proceedings of the 30th Conference on Uncertainty in Artificial Intelligence, in Proceedings of Machine Learning Research R12:17-26 Available from https://proceedings.mlr.press/r12/foulds14a.html. Reissued by PMLR on 04 October 2026.

Related Material