2015/07/16 by Amin Jalali, Maryam Fazel, Jalali, Amin +3 · 1 citation
Computer Science · Engineering · Mathematics · Medicine · #Artificial intelligence #Computer science #Conic optimization #Convex analysis #Convex optimization #Convexity #Discrete mathematics #Disjoint sets #FOS: Computer and information sciences #FOS: Mathematics #Kernel (algebra) #Kernel method #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mathematical optimization #Mathematics #Optimization and Control (math.OC) #Optimization and Variational Analysis #Optimization problem #Orthogonality #Proper convex function #Radial basis function kernel #Regular polygon #Representer theorem #Simple (philosophy) #Sparse and Compressive Sensing Techniques #Subderivative #Support vector machine #Systemic Lupus Erythematosus Research #cs.LG #math.OC #stat.ML
paper · pdf · doi:10.48550/arxiv.1507.04734
26 pages, 5 figures, additional revisions to text, under revision in SIOPT, An earlier version of this work has appeared as Chapter 3 in reference [21]
openalex publication_date 2015/07/16 · arxiv created 2017/04/12 · arxiv updated 2017/04/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We propose a new class of convex penalty functions, called variational Gram functions (VGFs), that can promote pairwise relations, such as orthogonality, among a set of vectors in a vector space. These functions can serve as regularizers in convex optimization problems arising from hierarchical classification, multitask learning, and estimating vectors with disjoint supports, among other applications. We study convexity for VGFs, and give efficient characterizations for their convex conjugates, subdifferentials, and proximal operators. We discuss efficient optimization algorithms for regularized loss minimization problems where the loss admits a common, yet simple, variational representation and the regularizer is a VGF. These algorithms enjoy a simple kernel trick, an efficient line search, as well as computational advantages over first order methods based on the subdifferential or proximal maps. We also establish a general representer theorem for such learning problems. Lastly, numerical experiments on a hierarchical classification problem are presented to demonstrate the effectiveness of VGFs and the associated optimization algorithms.