2011/10/10 by Eduardo F. Mendes, Eduardo Mendes, Mendes, Eduardo F. +2 · 1 citation
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (stat.ML) #Methodology (stat.ME) #Statistical Methods and Bayesian Inference #Statistical Methods and Inference #Statistics Theory (math.ST) #math.ST #stat.ME #stat.ML #stat.TH
paper · pdf · doi:10.48550/arxiv.1110.2058
openalex publication_date 2011/10/10 · arxiv created 2011/11/01 · arxiv updated 2011/11/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In mixtures-of-experts (ME) model, where a number of submodels (experts) are combined, there have been two longstanding problems: (i) how many experts should be chosen, given the size of the training data? (ii) given the total number of parameters, is it better to use a few very complex experts, or is it better to combine many simple experts? In this paper, we try to provide some insights to these problems through a theoretic study on a ME structure where m experts are mixed, with each expert being related to a polynomial regression model of order k. We study the convergence rate of the maximum likelihood estimator (MLE), in terms of how fast the Kullback-Leibler divergence of the estimated density converges to the true density, when the sample size n increases. The convergence rate is found to be dependent on both m and k, and certain choices of m and k are found to produce optimal convergence rates. Therefore, these results shed light on the two aforementioned important problems: on how to choose m, and on how m and k should be compromised, for achieving good convergence rates.