2018/04/10 by Igor Fedorov, Bhaskar D. Rao, Fedorov, Igor +1
Computer Science · Engineering · #Blind Source Separation Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Remote-Sensing Image Classification #Text and Document Classification Technologies
paper · pdf · doi:10.48550/arxiv.1804.03740
openalex publication_date 2018/04/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This paper addresses the problem of learning dictionaries for multimodal datasets, i.e. datasets collected from multiple data sources. We present an algorithm called multimodal sparse Bayesian dictionary learning (MSBDL). MSBDL leverages information from all available data modalities through a joint sparsity constraint. The underlying framework offers a considerable amount of flexibility to practitioners and addresses many of the shortcomings of existing multimodal dictionary learning approaches. In particular, the procedure includes the automatic tuning of hyperparameters and is unique in that it allows the dictionaries for each data modality to have different cardinality, a significant feature in cases when the dimensionality of data differs across modalities. MSBDL is scalable and can be used in supervised learning settings. Theoretical results relating to the convergence of MSBDL are presented and the numerical results provide evidence of the superior performance of MSBDL on synthetic and real datasets compared to existing methods.