Stable Deep Reinforcement Learning via Isotropic Gaussian Representations
2026/02/22 by Pasand, Ali Saheb, Ali Saheb Pasand, Johan Obando-Ceron +3 · 1 voice
Computer Science · #Domain Adaptation and Few-Shot Learning #Entropy (arrow of time) #Gaussian #Gaussian Processes and Bayesian Inference #Gaussian process #Isotropy #Regularization (linguistics) #Reinforcement Learning in Robotics #Reinforcement learning #Representation (politics) #Training set #cs.AI #cs.LG
paper · pdf · doi:10.48550/arxiv.2602.19373
openalex publication_date 2026/02/22 · arxiv published 2026/02/22 · openalex created_date 2026/02/26 · arxiv updated 2026/06/04 · openalex updated_date 2026/07/28
Abstract
Deep reinforcement learning systems often suffer from unstable training dynamics due to non-stationarity, where learning objectives and data distributions evolve over time. We show that under non-stationary targets, isotropic Gaussian embeddings are provably advantageous. In particular, they induce stable tracking of time-varying targets for linear readouts, achieve maximal entropy under a fixed variance budget, and encourage a balanced use of all representational dimensions--all of which enable agents to be more adaptive and stable. Building on this insight, we propose the use of Sketched Isotropic Gaussian Regularization for shaping representations toward an isotropic Gaussian distribution during training. We demonstrate empirically, over a variety of domains, that this simple and computationally inexpensive method improves performance under non-stationarity while reducing representation collapse, neuron dormancy, and training instability.
Citations
- Chimère Ω — blueprint for a physico-cognitively inspired local-first LLM runtime
- Simplicial Embeddings Improve Sample Efficiency in Actor-Critic Agents
- Use the Online Network If You Can: Towards Fast and Stable Reinforcement Learning
- A Survey of State Representation Learning for Deep Reinforcement Learning
- Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning
- The Courage to Stop: Overcoming Sunk Cost Fallacy in Deep Reinforcement Learning
- The Impact of On-Policy Parallelized Data Collection on Deep Reinforcement Learning Networks
- Mitigating Plasticity Loss in Continual Reinforcement Learning by Reducing Churn
- Measure gradients, not activations! Enhancing neuronal activity in deep reinforcement learning
- Meta-World+: An Improved, Standardized, RL Benchmark
- Towards General-Purpose Model-Free Reinforcement Learning
- SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement Learning
- Don't flatten, tokenize! Unlocking the key to SoftMoE's efficacy in deep RL
- Simplifying Deep Temporal Difference Learning
- Normalization and effective learning rates in reinforcement learning
- Mixture of Experts in a Mixture of RL settings
- On the consistency of hyper-parameter selection in value-based deep reinforcement learning
- No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
- In value-based deep reinforcement learning, a pruned network is a good network
- Mixtures of Experts Unlock Parameter Scaling for Deep RL
- Understanding and Addressing the Pitfalls of Bisimulation-based Representations in Offline Reinforcement Learning
- Small batch deep reinforcement learning
- For SALE: State-Action Representation Learning for Deep Reinforcement Learning
- Bigger, Better, Faster: Human-level Atari with human-level efficiency
- Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks
- Understanding plasticity in neural networks
- The Dormant Neuron Phenomenon in Deep Reinforcement Learning
- Atari-5: Distilling the Arcade Learning Environment down to Five Games
- Hyperbolic Deep Reinforcement Learning
- The Primacy Bias in Deep Reinforcement Learning
- Understanding and Preventing Capacity Loss in Reinforcement Learning
- Deep Reinforcement Learning at the Edge of the Statistical Precipice
- Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning
- VICReg: Variance-Invariance-Covariance Regularization for\n Self-Supervised Learning
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement Learning
- Data-Efficient Reinforcement Learning with Self-Predictive Representations
- Learning Invariant Representations for Reinforcement Learning without Reconstruction
- Transient Non-Stationarity and Generalisation in Deep Reinforcement Learning
- Scalable methods for computing state similarity in deterministic Markov\n Decision Processes
- On the Variance of the Adaptive Learning Rate and Beyond
- DeepMDP: Learning Continuous Latent Space Models for Representation Learning
- Deep Reinforcement Learning and the Deadly Triad
- Addressing Function Approximation Error in Actor-Critic Methods
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Proximal Policy Optimization Algorithms
- Optimizing Neural Networks with Kronecker-factored Approximate Curvature
- Human-level control through deep reinforcement learning
- The Arcade Learning Environment: An Evaluation Platform for General Agents
- Methods for computing state similarity in Markov Decision Processes
- Metrics for labelled Markov processes
- Bridging the Performance Gap Between Target-Free and Target-Based Reinforcement Learning
- Reinforcement Learning: An Introduction
Discussions
Related