Model-Based Clustering, Discriminant Analysis, and Density Estimation
2002/06/01 by Chris Fraley, Adrian E. Raftery, Adrian E Raftery · 150 citations
Computer Science · Mathematics · #Bayesian Methods and Mixture Models #Advanced Clustering Algorithms Research #Statistical Methods and Inference
paper · doi:10.1198/016214502760047131
Abstract
Cluster analysis is the automated search for groups of related observations in a dataset. Most clustering done in practice is based largely on heuristic but intuitively reasonable procedures, and most clustering methods available in commercial software are also of this type. However, there is little systematic guidance associated with these methods for solving important practical questions that arise in cluster analysis, such as how many clusters are there, which clustering method should be used, and how should outliers be handled. We review a general methodology for model-based clustering that provides a principled statistical approach to these issues. We also show that this can be useful for other problems in multivariate analysis, such as discriminant analysis and multivariate density estimation. We give examples from medical diagnosis, minefield detection, cluster recovery from noisy data, and spatial density estimation. Finally, we mention limitations of the methodology and discuss recent developments in model-based clustering for non-Gaussian data, high-dimensional datasets, large datasets, and Bayesian estimation.
Citations
Cited by
- cardinalR: Generating Interesting High-Dimensional Data Structures
- A flexible class of latent variable models for the analysis of antibody response data
- Parsimonious Ultrametric Manly Mixture Models
- A latent variable model for identifying and characterizing food adulteration
- Semiparametric Robust Estimation of Population Location
- SMLSOM: The shrinking maximum likelihood self-organizing map
- Benchmarking of Clustering Validity Measures Revisited
- A Bayesian Mark Interaction Model for Analysis of Tumor Pathology Images
- Bayesian nonparametric location-scale-shape mixtures
- K-Means and Gaussian Mixture Modeling with a Separation Constraint
- Clustering Airbnb Reviews
- Survey of Clustering Algorithms
- The Cartography of Opportunity: Spatial Data Science for Equitable Urban Policy
- Unconstrained representation of orthogonal matrices with application to common principle components
- Unobserved classes and extra variables in high-dimensional discriminant analysis
- Efficient Learning for Clustering and Optimizing Context-Dependent Designs
- A local depth measure for general data
- A tutorial on spectral clustering
- 5-Hydroxymethylcytosine signatures in cell-free DNA provide information about tumor types and stages
- Modeling and Predicting Power Consumption of High Performance Computing Jobs
- A 15,800 year record of productivity, carbon accumulation and environmental change at Eight Mile Lake and its catchment, central Alaska
- Kernel discriminant analysis and clustering with parsimonious Gaussian process models
- Quantile-based classifiers
- Integrating Fuzzy Logic and Statistics to Improve the Reliable Delimitation of Biogeographic Regions and Transition Zones
- A concave pairwise fusion approach to subgroup analysis
- Deep Gaussian Mixture Models
- Inferring alternative ecosystem states with field survey data
- Relating latent class membership to covariates and outcomes: Two bias-adjusted methods in Stata
- Deep Gaussian mixture models
- Mixtures Closest to a Given Measure: A Semidefinite Programming Approach
- Traits Without Borders: Integrating Functional Diversity Across Scales
- Graph-based Learning with Unbalanced Clusters
- Temperature and CO2 interactively drive shifts in the compositional and functional structure of peatland protist communities
- Morphological and molecular evidence for major re-circumscriptions in and eight new species of Melichrus R.Br. (Ericaceae subfam. Epacridoideae) in eastern Australia
- Lethal conflict after group fission in wild chimpanzees
- A Model-based Semi-Supervised Clustering Methodology
- Hybridization of Expectation-Maximization and K-Means Algorithms for Better Clustering Performance
- Investigation of Parameter Uncertainty in Clustering Using a Gaussian Mixture Model Via Jackknife, Bootstrap and Weighted Likelihood Bootstrap
- Opportunities and Challenges in Applying AI to Evolutionary Morphology
- Effective Clustering for Large Multi-Relational Graphs
- A Doubly-Enhanced EM Algorithm for Model-Based Tensor Clustering
- A Parsimonious Explanation of the Resilient, Undercontrolled, and Overcontrolled Personality Types
- Exponential Series Approaches for Nonparametric Graphical Models
- Agglomerative and divisive hierarchical Bayesian clustering
- Robustification of Elliott's on-line EM algorithm for HMMs
- Dirichlet Process Parsimonious Mixtures for clustering
- Seq2Karyotype (S2K): A Method for in-silico Karyotyping Using Single-Sample Whole-Genome Sequencing Data
- Fitting A Mixture Distribution to Data: Tutorial
- Latent Simplex Position Model: High Dimensional Multi-view Clustering with Uncertainty Quantification
- Spatial Clustering of Time-Series via Mixture of Autoregressions Models and Markov Random Fields
- Maximum likelihood estimation of a multidimensional log-concave density
- Diagonally-Weighted Generalized Method of Moments Estimation for Gaussian Mixture Modeling
- Multivariate normal mixture modeling, clustering and classification with the rebmix package
- Variable selection for clustering with Gaussian mixture models: state of the art
- Graph Construction for Learning with Unbalanced Data
- Gaussian Mixture Models with Component Means Constrained in Pre-selected Subspaces
- How many clusters? An information theoretic perspective
- Subgroup Analysis for Longitudinal data via Semiparametric Additive Mixed Effect Model
- Regularized k-POD: Sparse k-means clustering for high-dimensional missing data
- Adaptive Clustering through Semidefinite Programming
- Identifiability of Hierarchical Latent Attribute Models
- Block clustering with collapsed latent block models
- Modelling overdispersion heterogeneity in differential expression analysis using mixtures
- A Family of Mixture Models for Biclustering
- Handling missing data in model-based clustering
- Normalizing Flow to Augmented Posterior: Conditional Density Estimation with Interpretable Dimension Reduction for High Dimensional Data
- Combinatorial clustering and the beta negative binomial process
- Parsimonious Gaussian mixture models with piecewise-constant eigenvalue profiles
- Mixture model modal clustering
- A Dirichlet Process Mixture Model for Clustering Longitudinal Gene Expression Data
- Exploring the space-time pattern of log-transformed infectious count of COVID-19: a clustering-segmented autoregressive sigmoid model
- Mixture Envelope Model for Heterogeneous Genomics Data Analysis
- Partial membership models for soft clustering of multivariate football player performance data
- Modifications of the BIC for order selection in finite mixture models
- Sparse Bayesian Hierarchical Modeling of High-dimensional Clustering Problems
- Mixture model with multiple allocations for clustering spatially correlated observations in the analysis of ChIP-Seq data
- One-class classification with application to forensic analysis
- Identifiability and consistency of network inference using the hub model and variants
- Simultaneous Dimensionality and Complexity Model Selection for Spectral Graph Clustering
- Machine Unlearning for Robust DNNs: Attribution-Guided Partitioning and Neuron Pruning in Noisy Environments
- Mixtures of Multivariate Power Exponential Distributions
- Identifying the number of clusters in discrete mixture models
- Distributed Learning of Finite Gaussian Mixtures
- Estimation and Model Selection for Model-Based Clustering with the Conditional Classification Likelihood
- A model selection approach for clustering a multinomial sequence with non-negative factorization
- On the Optimality of Kernel-Embedding Based Goodness-of-Fit Tests
- Semi-supervised classification for dynamic Android malware detection
- Mixture regression for observational data, with application to functional regression models
- On the characterization of flowering curves using Gaussian mixture models
- The Importance of Being Correlated: Implications of Dependence in Joint Spectral Inference across Multiple Networks
- Bayesian nonparametric clustering for spatio-temporal data, with an application to air pollution
- Flexible parametric bootstrap for testing homogeneity against clustering and assessing the number of clusters
- Modeling and estimating skewed and heavy-tailed populations via unsupervised mixture models
- Model Based Clustering for Mixed Data: clustMD
- Spectral Clustering with Imbalanced Data
- Model-Based Clustering with Sequential Outlier Identification using the Distribution of Mahalanobis Distances
- Flexible High-Dimensional Unsupervised Learning with Missing Data
- On Bayesian Modelling of the Uncertainties in Palaeoclimate Reconstruction
- Selective Clustering Annotated using Modes of Projections
- mtDNA variation in East Africa unravels the history of afro‐asiatic groups
- Unsupervised learning of regression mixture models with unknown number of components
- Comparison of Clustering Methods for Time Course Genomic Data: Applications to Aging Effects
- Clusters and water flows: a novel approach to modal clustering through Morse theory
- P-values for classification
- Bounds for Bayesian order identification with application to mixtures
- Treelets—An adaptive multi-scale basis for sparse unordered data
- Pollen morphology and its taxonomic utility in the Southern Hemisphere bracteate‐prostrate forget‐me‐nots ( Myosotis , Boraginaceae)
- Morphological analyses support recognition of three new threatened species of bracteate–prostrate Myosotis (Boraginaceae) endemic to the South Island of Aotearoa New Zealand
- Improved Density-Based Spatio--Textual Clustering on Social Media
- Bayesian model selection for the latent position cluster model for Social Networks
- A robust method for cluster analysis
- Cutoff for exact recovery of Gaussian mixture models
- Species of Dickinsonia Sprigg from the Ediacaran of South Australia
- Multidimensional data-driven classification of emission-line galaxies
- Data Stream Clustering: Challenges and Issues
- Evidence of an Upper Bound on the Masses of Planets and Its Implications for Giant Planet Formation
- Linguistic, geographic and genetic isolation: a collaborative study of Italian populations.
- Widespread ancient whole‐genome duplications in Malpighiales coincide with Eocene global climatic upheaval
- Massive migration from the steppe is a source for Indo-European languages in Europe
- A Hybrid Mixture of t-Factor Analyzers for Clustering High-dimensional Data
- Sparse mixed linear modeling with anchor-based guidance for high-entropy alloy discovery
- An Admixture Approach to Trihybrid Ancestry Variation in the Philippines with Implications for Forensic Anthropology
- Empirical Bayes, SURE and Sparse Normal Mean Models
- Mean-field theory of Bayesian clustering
- Treatment of bimodality in proficiency test of pH in bioethanol matrix
- Nonparametric Unsupervised Classification
- Shared language, diverging genetic histories: high-resolution analysis of Y-chromosome variability in Calabrian and Sicilian Arbereshe
- Species limits and taxonomic revision of the bracteate-prostrate group of southern hemisphere forget-me-nots (Myosotis, Boraginaceae), including description of three new species endemic to New Zealand
- Evolution and speciation in the Eocene planktonic foraminifer Turborotalia
- Latent Class Detection and Class Assignment: A Comparison of the MAXEIG Taxometric Procedure and Factor Mixture Modeling Approaches
- Solution path clustering with adaptive concave penalty
- Functional clustering in nested designs: Modeling variability in reproductive epidemiology studies
- Leveraging Noisy Manual Labels as Useful Information: An Information Fusion Approach for Enhanced Variable Selection in Penalized Logistic Regression
- Young star clusters in nearby molecular clouds
- Learning the smoothness of noisy curves with application to online curve estimation
- Learning mixtures of Bernoulli templates by two-round EM with performance guarantee
- Order preserving hierarchical agglomerative clustering
- An Ancient Mediterranean Melting Pot: Investigating the Uniparental Genetic Structure and Population History of Sicily and Southern Italy
- Model-based clustering in very high dimensions via adaptive projections
- A biclustering approach to university performances: an Italian case\n study
- On a two-truths phenomenon in spectral graph clustering
- THE SPECTRAL ENERGY DISTRIBUTIONS OF FERMI BLAZARS
- Sloan Great Wall as a complex of superclusters with collapsing cores
- Model-based SIR for dimension reduction
- Multi-layered characterisation of hot stellar systems with confidence
- Anomaly and Novelty detection for robust semi-supervised learning
- A Novel Framework Using Unsupervised Machine Learning for Aerosol Classification: A Case Study of North Indian AERONET Sites
- Optimal Bayesian clustering using non-negative matrix factorization
- Statistical few-shot learning for large-scale classification via parameter pooling
- Machine learning in epilepsy
Related