2013/03/29 by Guido Montúfar, Guido F. Montúfar, Montúfar, Guido F.
Computer Science · Mathematics · #60C05 #68Q32 #82C32 #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Probability (math.PR) #Stochastic Gradient Optimization Techniques #cs.LG #math.PR #msc:60C05 #msc:68Q32 #msc:82C32 #stat.ML
paper · pdf · doi:10.48550/arxiv.1303.7461
19 pages, 5 figures, 1 table
openalex publication_date 2013/03/29 · arxiv created 2014/01/28 · arxiv updated 2014/01/30 · openalex created_date 2022/10/04 · openalex updated_date 2026/07/28
We generalize recent theoretical work on the minimal number of layers of narrow deep belief networks that can approximate any probability distribution on the states of their visible units arbitrarily well. We relax the setting of binary units (Sutskever and Hinton, 2008; Le Roux and Bengio, 2008, 2010; Montúfar and Ay, 2011) to units with arbitrary finite state spaces, and the vanishing approximation error to an arbitrary approximation error tolerance. For example, we show that a q-ary deep belief network with L≥ 2+\fracq\lceil m-δ\rceil-1q-1 layers of width n ≤ m + logq(m) + 1 for some m∈ ℕ can approximate any probability distribution on \0,1,…,q-1\n without exceeding a Kullback-Leibler divergence of δ. Our analysis covers discrete restricted Boltzmann machines and naïve Bayes models as special cases.