2025/06/03 by Luca Arnaboldi, Arnaboldi, Luca, Bruno Loureiro +7 · 3 citations
Computer Science · #Class (philosophy) #Convergence (economics) #Disordered Systems and Neural Networks (cond-mat.dis-nn) #Encoding (memory) #FOS: Computer and information sciences #FOS: Physical sciences #Generative Adversarial Networks and Image Synthesis #Initialization #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #Population #Sequence (biology) #Sequence learning #Space (punctuation) #Stochastic Gradient Optimization Techniques
paper · pdf · doi:10.48550/arxiv.2506.02651
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/06/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We study the dynamics of stochastic gradient descent (SGD) for a class of sequence models termed Sequence Single-Index (SSI) models, where the target depends on a single direction in input space applied to a sequence of tokens. This setting generalizes classical single-index models to the sequential domain, encompassing simplified one-layer attention architectures. We derive a closed-form expression for the population loss in terms of a pair of sufficient statistics capturing semantic and positional alignment, and characterize the induced high-dimensional SGD dynamics for these coordinates. Our analysis reveals two distinct training phases: escape from uninformative initialization and alignment with the target subspace, and demonstrates how the sequence length and positional encoding influence convergence speed and learning trajectories. These results provide a rigorous and interpretable foundation for understanding how sequential structure in data can be beneficial for learning with attention-based models.