vix.ing · top · new · best · stats · spec

Mahankali, Arvind

  1. One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention
    2023/07/07 by Arvind V. Mahankali, Tatsunori Hashimoto, Mahankali, Arvind +3 · 31 citations
    Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
  2. Beyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and Time
    2023/06/28 by Mahankali, Arvind, Haochen, Jeff Z., Dong, Kefan +2 · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG)