Transformers learn in-context by gradient descent
2022/12/15 by Johannes von Oswald, Eyvind Niklasson, von Oswald, Johannes +11 · 6 voices · 196 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Artificial intelligence #Artificial neural network #Computer science #Domain Adaptation and Few-Shot Learning #Engineering #Generative Adversarial Networks and Image Synthesis #Gradient descent #Machine learning #Stochastic gradient descent #Transformer #cs.AI #cs.CL #cs.LG
paper · pdf · doi:10.48550/arxiv.2212.07677
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2022/12/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
Abstract
At present, the mechanisms of in-context learning in Transformers are not well understood and remain mostly an intuition. In this paper, we suggest that training Transformers on auto-regressive objectives is closely related to gradient-based meta-learning formulations. We start by providing a simple weight construction that shows the equivalence of data transformations induced by 1) a single linear self-attention layer and by 2) gradient-descent (GD) on a regression loss. Motivated by that construction, we show empirically that when training self-attention-only Transformers on simple regression tasks either the models learned by GD and Transformers show great similarity or, remarkably, the weights found by optimization match the construction. Thus we show how trained Transformers become mesa-optimizers i.e. learn models by gradient descent in their forward pass. This allows us, at least in the domain of regression problems, to mechanistically understand the inner workings of in-context learning in optimized Transformers. Building on this insight, we furthermore identify how Transformers surpass the performance of plain gradient descent by learning an iterative curvature correction and learn linear models on deep data representations to solve non-linear regression tasks. Finally, we discuss intriguing parallels to a mechanism identified to be crucial for in-context learning termed induction-head (Olsson et al., 2022) and show how it could be understood as a specific case of in-context learning by gradient descent learning within Transformers. Code to reproduce the experiments can be found at https://github.com/google-research/self-organising-systems/tree/master/transformerslearniclbygd .
Cited by
- Tabular Foundation Models for Discrete Choice Estimation
- Vector Search As Nearest Neighbor Matching: RAG-based Policy Learning in Causal Inference
- Topological Signatures of Context-Level Reliability in TabPFN
- Bigger Is Safer: Provable Robustness in In-Context Learning Scales with Capacity
- Loop the Loopies!
- What does a Bayes-filtered transformer believe? A predictive Monte Carlo approach
- In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention
- A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems
- Bayesian Wind Tunnels for Model Selection
- Lifted State Hypothesis in Large Language Models
- Transformers are Bayesian Networks
- Transformers for dynamical systems learn transfer operators in-context
- Remapping and navigation of an embedding space via error minimization: a fundamental organizational principle of cognition in natural and artificial systems
- Shared sensitivity to data distribution during learning in humans and transformer networks
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
- How Do LLMs Use Their Depth?
- Fast weight programming and linear transformers: from machine learning to neurobiology
- Learning without training: The implicit dynamics of in-context learning
- What Neuroscience Can Teach AI About Learning in Continuously Changing Environments
- Relational reasoning and inductive bias in transformers and large language models
- True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics
- The Geometry of Prompting: Unveiling Distinct Mechanisms of Task Adaptation in Language Models
- Training Dynamics of In-Context Learning in Linear Attention
- To Use or not to Use Muon: How Simplicity Bias in Optimizers Matters
- Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners
- In-Context Learning as Implicit Policy Gradient
- MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models
- Task Schema and Binding: A Double Dissociation Study of In-Context Learning
- In-Context Algebra
- NRGPT: An Energy-based Alternative for GPT
- In-Context Multi-Operator Learning with DeepOSets
- In-Context Semi-Supervised Learning
- Softmax as Linear Attention in the Large-Prompt Regime: a Measure-based Perspective
- LAPPI: Interactive Optimization with LLM-Assisted Preference-Based Problem Instantiation
- Rethinking Label Consistency of In-Context Learning: An Implicit Transductive Label Propagation Perspective
- Unlocking the Address Book: Dissecting the Sparse Semantic Structure of LLM Key-Value Caches via Sparse Autoencoders
- LogICL: Distilling LLM Reasoning to Bridge the Semantic Gap in Cross-Domain Log Anomaly Detection
- Impact of Positional Encoding: Clean and Adversarial Rademacher Complexity for Transformers under In-Context Regression
- Supervised learning pays attention
- In-Context and Few-Shots Learning for Forecasting Time Series Data based on Large Language Models
- The Initialization Determines Whether In-Context Learning Is Gradient Descent
- The brain-AI convergence: Predictive and generative world models for general-purpose computation
- Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
- Leveraging Language Models for Interpretable Analysis of Narratives in a Large Corpus
- Equivalence of Context and Parameter Updates in Modern Transformer Blocks
- Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
- The Specification Trap: Why Static Value Alignment Alone Is Insufficient for Robust Alignment
- Optimal Attention Temperature Improves the Robustness of In-Context Learning under Distribution Shift in High Dimensions
- Convergence dynamics of Agent-to-Agent Interactions with Misaligned objectives
- Misaligned by Design: Incentive Failures in Machine Learning
- Implicit Federated In-context Learning For Task-Specific LLM Fine-Tuning
- Scaling Laws and In-Context Learning: A Unified Theoretical Framework
- On the Emergence of Induction Heads for In-Context Learning
- Detecting Data Contamination in LLMs via In-Context Learning
- How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
- Take Goodhart Seriously: Principled Limit on General-Purpose AI Optimization
- Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime
- The Consensus Trap: Rescuing Multi-Agent LLMs from Adversarial Majorities via Token-Level Collaboration
- On Surprising Effectiveness of Masking Updates in Adaptive Optimizers
- Context Structure Reshapes the Representational Geometry of Language Models
- The Bayesian Geometry of Transformer Attention
- Understanding Multi-View Transformers
- Provable test-time adaptivity and distributional robustness of in-context learning
- Can Language Models Compose Skills In-Context?
- A Framework for Quantifying How Pre-Training and Context Benefit In-Context Learning
- Enabling Robust In-Context Memory and Rapid Task Adaptation in Transformers with Hebbian and Gradient-Based Plasticity
- Large Language Models as Model Organisms for Human Associative Learning
- Do Prompts Reshape Representations? An Empirical Study of Prompting Effects on Embeddings
- Transformers are almost optimal metalearners for linear classification
- Optimality and NP-Hardness of Transformers in Learning Markovian Dynamical Functions
- Layer Specialization Underlying Compositional Reasoning in Transformers
- Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
- Rewiring Experts on the Fly:Continuous Rerouting for Better Online Adaptation in Mixture-of-Expert models
- Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
- Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
- Dimension-Free Minimax Rates for Learning Pairwise Interactions in Attention-Style Models
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
- Softmax ≥ Linear: Transformers may learn to classify in-context by kernel gradient descent
- Compositional meta-learning through probabilistic task inference
- Design Principles for Sequence Models via Coefficient Dynamics
- On the Relationship Between the Choice of Representation and In-Context Learning
- Transmuting prompts into weights
- Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models
- Fine-Grained Emotion Recognition via In-Context Learning
- ContextNav: Towards Agentic Multimodal In-Context Learning
- Learning Linear Regression with Low-Rank Tasks in-Context
- Multi-Agent Collaborative Intelligence: Dual-Dial Control for Reliable LLM Reasoning
- Allocation of Parameters in Transformers
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and Time
- Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
- Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
- Pretrain-Test Task Alignment Governs Generalization in In-Context Learning
- Test time training enhances in-context learning of nonlinear functions
- TTT3R: 3D Reconstruction as Test-Time Training
- ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
- In-Context Compositional Q-Learning for Offline Reinforcement Learning
- From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs
- Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention
- IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
- Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
- In-Context Learning can Perform Continual Learning Like Humans
- A circuit for predicting hierarchical structure in-context in Large Language Models
- On Theoretical Interpretations of Concept-Based In-Context Learning
- Linear Transformers Implicitly Discover Unified Numerical Algorithms
- Verbalizing LLM's Higher-order Uncertainty via Imprecise Probabilities
- Asymptotic Study of In-context Learning with Random Transformers through Equivalent Models
- Selective Induction Heads: How Transformers Select Causal Structures In Context
- Just-in-time and distributed task representations in language models
- InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity
- Learning In-context \pmbn-grams with Transformers: Sub-\pmbn-grams Are Near-stationary Points
- Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
- Provable Low-Frequency Bias of In-Context Learning of Representations
- In-Context Reinforcement Learning via Communicative World Models
- From Text to Trajectories: GPT-2 as an ODE Solver via In-Context
- Understanding Transformers through the Lens of Pavlovian Conditioning
- Transformers in Pseudo-Random Number Generation: A Dual Perspective on Theory and Practice
- Provable In-Context Learning of Nonlinear Regression with Transformers
- Towards Compute-Optimal Many-Shot In-Context Learning
- Rethinking Invariance in In-context Learning
- On Finetuning Tabular Foundation Models
- LLMs are Bayesian, In Expectation, Not in Realization
- Transformers Don't In-Context Learn Least Squares Regression
- Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations
- Asymptotic theory of in-context learning by linear attention
- An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
- Decomposing Prediction Mechanisms for In-Context Recall
- ICLShield: Exploring and Mitigating In-Context Learning Backdoor Attacks
- Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
- CooT: Learning to Coordinate In-Context with Coordination Transformers
- Can Gradient Descent Simulate Prompting?
- In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly
- Finding Clustering Algorithms in the Transformer Architecture
- CausalPFN: Amortized Causal Effect Estimation via In-Context Learning
- Federated In-Context Learning: Iterative Refinement for Improved Answer Quality
- Latent Concept Disentanglement in Transformer-based Language Models
- Next-Token Prediction Should be Ambiguity-Sensitive: A Meta-Learning Perspective
- When and How Unlabeled Data Provably Improve In-Context Learning
- Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
- Can Biologically Plausible Temporal Credit Assignment Rules Match BPTT for Neural Similarity? E-prop as an Example
- Understanding In-Context Learning on Structured Manifolds: Bridging Attention to Kernel Methods
- Understanding In-context Learning of Addition via Activation Subspaces
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
- Contextually Guided Transformers via Low-Rank Adaptation
- Sample Complexity and Representation Ability of Test-time Scaling Paradigms
- Transformers Meet In-Context Learning: A Universal Approximation Theory
- Counterfactual reasoning: an analysis of in-context emergence
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
- When can in-context learning generalize out of task distribution?
- ConText: Driving In-context Learning for Text Removal and Segmentation
- The Guanyin Protocol: A Framework for Immediately Establishing an Understanding of Both Causality and Compassion in LLM Systems Using Semantic Anchoring
- Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models
- Weight-Space Linear Recurrent Neural Networks
- From Parameters to Prompts: Understanding and Mitigating the Factuality Gap between Fine-Tuned LLMs
- Neither Stochastic Parroting nor AGI: LLMs Solve Tasks through Context-Directed Extrapolation from Training Data Priors
- The Role of Diversity in In-Context Learning for Large Language Models
- Optimization-Inspired Few-Shot Adaptation for Large Language Models
- Next-token pretraining implies in-context learning
- Understanding Prompt Tuning and In-Context Learning via Meta-Learning
- From Compression to Expression: A Layerwise Analysis of In-Context Learning
- Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
- Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse
- Joint Flashback Adaptation for Forgetting-Resistant Instruction Tuning
- Meta-Learning an In-Context Transformer Model of Human Higher Visual Cortex
- The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
- Why Large Language Models Fail at Tabular Prediction
- Out-of-Distribution Generalization of In-Context Learning: A Low-Dimensional Subspace Perspective
- Adversarially Pretrained Transformers may be Universally Robust In-Context Learners
- Attention-based clustering
- Transformer learns the cross-task prior and regularization for in-context learning
- Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform
- Do different prompting methods yield a common task representation in language models?
- Transformers as Unsupervised Learning Algorithms: A study on Gaussian Mixtures
- Context parroting: A simple but tough-to-beat baseline for foundation models in scientific machine learning
- Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
- Permutation Randomization on Nonsmooth Nonconvex Optimization: A Theoretical and Experimental Study
- An evolutionary perspective on modes of learning in Transformers
- Scaling In-Context Online Learning Capability of LLMs via Cross-Episode Meta-RL
- When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs
- Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
- On the generalization of language models from in-context learning and finetuning: a controlled study
- Training-Free Looped Transformers
- Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning
- Large language models reorganize representational geometry during in-context learning
- Universal priors: solving empirical Bayes via Bayesian inference and pretraining
- Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
- In-Context Learning can distort the relationship between sequence likelihoods and biological fitness
- Decoding Recommendation Behaviors of In-Context Learning LLMs Through Gradient Descent
- Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
- From predictions to confidence intervals: an empirical study of conformal prediction methods for in-context learning
- Scaling sparse feature circuit finding for in-context learning
- Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
- Can Pre-training Indicators Reliably Predict Fine-tuning Outcomes of LLMs?
- Align-RAG: Alignment Is All You Need for TSFM In-Context Learning
- Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
- When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
- How new data permeates LLM knowledge and how to dilute it
- Long Context In-Context Compression by Getting to the Gist of Gisting
- Multihead self-attention in cortico-thalamic circuits
Discussions
- 1. What Can Transformers Learn In-Context? A Case Study of Simple Function Classes
arxiv.org/abs/2208.01066
2. What learning algorithm is in-context learning? Investigations with linear models
arxiv.o [bsky, 14 points, 1 comments]
- My intuition (potentially wrong) is that ICL is probably analogous in some ways to more conventional learning algorithms, although I don’t feel like I totally understand in a rigorous way why it’s so [bsky, 2 points, 0 comments]
- Transformers learn in-context by gradient descent [hn, 2 points, 0 comments]
- Pretty sure this is the paper I had in mind (almost three years old now, so surely a lot has also built on it already as well) arxiv.org/abs/2212.07677 [bsky, 1 points, 1 comments]
- human learning is the same thing as ultra long context and context compression. in-context learning is learning and retrieval, synthesis, and so on (memory). I recommend reading von Neumann's The Comp [bsky, 0 points, 0 comments]
- 🔷 Link zum Fachbericht vom 31.05.2023 bei arXiv:
(Der Fachbericht ist mit einem Seiten-Translator übersetzbar.)
«Transformers learn in-context by gradient descent»
arxiv.org/abs/2212.076... [bsky, 0 points, 1 comments]
Related