2025/09/15 by Haoze He, Artemis Pados, He, Haoze +4
Computer Science · Engineering · Mathematics · #Cluster analysis #Convex optimization #Eigenvalues and eigenvectors #FOS: Electrical engineering #FOS: Mathematics #Face and Expression Recognition #Minimax #Numerical Analysis (math.NA) #Pattern recognition (psychology) #Selection (genetic algorithm) #Signal Processing (eess.SP) #Spectral clustering #Subspace topology #cs.NA #eess.SP #electronic engineering #information engineering #math.NA
paper · pdf · doi:10.48550/arxiv.2509.11981
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/09/15 · openalex created_date 2025/10/12 · arxiv created 2026/08/05 · openalex updated_date 2026/08/05 · arxiv updated 2026/08/07
We revisit the problem of spectral clustering in multimodal settings, where each data modality is encoded as a graph Laplacian. While classical approaches--including joint diagonalization, spectral co-regularization, and multiview clustering--attempt to align embeddings across modalities, they often rely on costly iterative refinement and may fail to directly target the spectral subspace relevant for clustering. In this work, we introduce two key innovations. First, we bring the power of randomization to this setting by sampling random convex combinations of Laplacians as a simple and scalable alternative to explicit eigenspace alignment. Second, we propose a principled selection rule based on Bottom-k Aggregated Spectral Energy (BASE)--a k-dimensional extension of the directional smoothness objective from recent minimax formulations--which we uniquely apply as a selection mechanism rather than an optimization target. The result is Randomized Joint Diagonalization with BASE Selection (RJD-BASE), a method that is easily implementable, computationally efficient, aligned with the clustering objective, and grounded in decades of progress in standard eigensolvers. Through experiments on synthetic and real-world datasets, we show that RJD-BASE reliably selects high-quality embeddings, outperforming classical multimodal clustering methods at low computational cost.