2021/12/27 by Carles Domingo-Enrich, Domingo-Enrich, Carles
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and ELM #Neural Networks and Applications #Stochastic Gradient Optimization Techniques #cs.LG
paper · pdf · doi:10.48550/arxiv.2112.13867
arxiv created 2021/12/27 · openalex publication_date 2021/12/27 · arxiv updated 2021/12/30 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
We construct pairs of distributions μd, νd on ℝd such that the quantity |𝔼x ∼ μd [F(x)] - 𝔼x ∼ νd [F(x)]| decreases as Ω(1/d2) for some three-layer ReLU network F with polynomial width and weights, while declining exponentially in d if F is any two-layer network with polynomial weights. This shows that deep GAN discriminators are able to distinguish distributions that shallow discriminators cannot. Analogously, we build pairs of distributions μd, νd on ℝd such that |𝔼x ∼ μd [F(x)] - 𝔼x ∼ νd [F(x)]| decreases as Ω(1/(dlog d)) for two-layer ReLU networks with polynomial weights, while declining exponentially for bounded-norm functions in the associated RKHS. This confirms that feature learning is beneficial for discriminators. Our bounds are based on Fourier transforms.