2024/12/21 by Songtao Lu, Lu, Songtao, Yingdong Lu +3
Computer Science · #Neural Networks and Applications #Reinforcement Learning in Robotics #Data Stream Mining Techniques
paper · pdf · doi:10.48550/arxiv.2412.16683
We derive the system of differential equations for the gradient flow characterizing the training process of linear in-context learning in full generality. Next, we explore the geometric structure of the gradient flows in two instances, including identifying its invariants, optimum, and saddle points. This understanding allows us to quantify the behavior of the two gradient flows under the full generality of parameters and data.