2020/08/31 by Patrick Huembeli, Alexandre Dauphin · 58 citations
Computer Science · Materials Science · Physics and Astronomy · #Artificial neural network #Benchmark (surveying) #Convergence (economics) #Eigenvalues and eigenvectors #Function (biology) #Gradient descent #Hessian matrix #Machine Learning in Materials Science #Maxima and minima #Neural Networks and Reservoir Computing #Quantum Computing Algorithms and Architecture #Stability (learning theory) #quant-ph
paper · pdf · doi:10.1088/2058-9565/abdbc9
published in Quantum Science and Technology 6(2), 025011 (IOP Publishing)
openalex created_date 2020/08/10 · openalex publication_date 2021/01/14 · arxiv created 2021/03/02 · arxiv updated 2021/03/24 · openalex updated_date 2026/08/05
Abstract Machine learning techniques enhanced by noisy intermediate-scale quantum (NISQ) devices and especially variational quantum circuits (VQC) have recently attracted much interest and have already been benchmarked for certain problems. Inspired by classical deep learning, VQCs are trained by gradient descent methods which allow for efficient training over big parameter spaces. For NISQ sized circuits, such methods show good convergence. There are however still many open questions related to the convergence of the loss function and to the trainability of these circuits in situations of vanishing gradients. Furthermore, it is not clear how ‘good’ the minima are in terms of generalization and stability against perturbations of the data and there is, therefore, a need for tools to quantitatively study the convergence of the VQCs. In this work, we introduce a way to compute the Hessian of the loss function of VQCs and show how to characterize the loss landscape with it. The eigenvalues of the Hessian give information on the local curvature and we discuss how this information can be interpreted and compared to classical neural networks. We benchmark our results on several examples, starting with a simple analytic toy model to provide some intuition about the behaviour of the Hessian, then going to bigger circuits, and also train VQCs on data. Finally, we show how the Hessian can be used to adjust the learning rate for faster convergence during the training of variational circuits.