Mahankali, Arvind
- One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention
2023/07/07 by Arvind V. Mahankali, Tatsunori Hashimoto, Mahankali, Arvind +3 · 31 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification
- Beyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and Time
2023/06/28 by Mahankali, Arvind, Haochen, Jeff Z., Dong, Kefan +2 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG)