2014/12/20 by Daniel Jiwoong Im, Im, Daniel Jiwoong, Ethan Buchman +3
Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Model Reduction and Neural Networks #cs.LG
paper · pdf · doi:10.48550/arxiv.1412.6617
Nine pages including the reference page plus one page appendix. Appeared at ICLR2015 workshop track
openalex publication_date 2014/12/20 · arxiv created 2015/04/07 · arxiv updated 2015/04/09 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
Energy-based models are popular in machine learning due to the elegance of their formulation and their relationship to statistical physics. Among these, the Restricted Boltzmann Machine (RBM), and its staple training algorithm contrastive divergence (CD), have been the prototype for some recent advancements in the unsupervised training of deep neural networks. However, CD has limited theoretical motivation, and can in some cases produce undesirable behavior. Here, we investigate the performance of Minimum Probability Flow (MPF) learning for training RBMs. Unlike CD, with its focus on approximating an intractable partition function via Gibbs sampling, MPF proposes a tractable, consistent, objective function defined in terms of a Taylor expansion of the KL divergence with respect to sampling dynamics. Here we propose a more general form for the sampling dynamics in MPF, and explore the consequences of different choices for these dynamics for training RBMs. Experimental results show MPF outperforming CD for various RBM configurations.