2017/07/25 by Seyed Mehran Kazemi, Bahare Fatemi, Kazemi, Seyed Mehran +13 · 1 citation
Computer Science · #Bayesian Modeling and Causal Inference #Data Mining Algorithms and Applications #Data Visualization and Analytics #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
paper · pdf · doi:10.48550/arxiv.1707.07785
openalex publication_date 2017/07/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Relational probabilistic models have the challenge of aggregation, where one variable depends on a population of other variables. Consider the problem of predicting gender from movie ratings; this is challenging because the number of movies per user and users per movie can vary greatly. Surprisingly, aggregation is not well understood. In this paper, we show that existing relational models (implicitly or explicitly) either use simple numerical aggregators that lose great amounts of information, or correspond to naive Bayes, logistic regression, or noisy-OR that suffer from overconfidence. We propose new simple aggregators and simple modifications of existing models that empirically outperform the existing ones. The intuition we provide on different (existing or new) models and their shortcomings plus our empirical findings promise to form the foundation for future representations.