Pascanu, Razvan
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
2024/02/29 by Soham De, De, Soham, Samuel L. Smith +31 · 5 voices · 28 citations
#cs.LG #cs.CL
- Relational inductive biases, deep learning, and graph networks
2018/06/04 by Peter W. Battaglia, Jessica B. Hamrick, Battaglia, Peter W. +51 · 2 voices · 120 citations
#cs.LG #cs.AI #stat.ML
- Progressive Neural Networks
2016/06/15 by Andrei A. Rusu, Neil C. Rabinowitz, Rusu, Andrei A. +13 · 3 voices · 94 citations
Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Neural Networks and Reservoir Computing #Reinforcement Learning in Robotics #cs.LG
- A simple neural network module for relational reasoning
2017/06/05 by Adam Santoro, David Raposo, Santoro, Adam +11 · 1 voice · 17 citations
#cs.CL #cs.LG
- On the difficulty of training Recurrent Neural Networks
2012/11/21 by Razvan Pascanu, Tomáš Mikolov, Pascanu, Razvan +3 · 275 citations
Computer Science · Physics and Astronomy · #Neural Networks and Applications #Model Reduction and Neural Networks #Domain Adaptation and Few-Shot Learning
- NoProp: Training Neural Networks without Full Back-propagation or Full Forward-propagation
2025/03/31 by Qinyu Li, Yee Whye Teh, Li, Qinyu +3 · 13 voices · 5 citations
Computer Science · #Neural Networks and Applications
- Resurrecting Recurrent Neural Networks for Long Sequences
2023/03/11 by Antonio Orvieto, Samuel Smith, Samuel L Smith +13 · 2 voices · 58 citations
Computer Science · Engineering · Materials Science · #cs.LG
- Uncovering mesa-optimization algorithms in Transformers
2023/09/11 by Johannes von Oswald, von Oswald, Johannes, Maximilian Schlegel +24 · 2 voices · 13 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.LG
- RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
2024/04/11 by Aleksandar Botev, Soham De, Botev, Aleksandar +127 · 1 voice · 10 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech Recognition and Synthesis
- Sobolev Training for Neural Networks
2017/06/15 by Wojciech Marian Czarnecki, Czarnecki, Wojciech Marian, Simon Osindero +7 · 1 voice · 15 citations
Computer Science · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning and Data Classification #cs.LG
- Sharp Minima Can Generalize For Deep Nets
2017/03/15 by Laurent Dinh, Dinh, Laurent, Razvan Pascanu +5 · 55 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Stochastic Gradient Optimization Techniques
- On the Number of Linear Regions of Deep Neural Networks
2014/02/08 by Montúfar, Guido, Pascanu, Razvan, Cho, Kyunghyun +1 · 40 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- Policy Distillation
2015/11/19 by Andrei A. Rusu, Sergio Gómez Colmenarejo, Rusu, Andrei A. +15 · 44 citations
Computer Science · Engineering · #Neural Networks and Reservoir Computing #CCD and CMOS Imaging Sensors #Visual Attention and Saliency Detection
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
2014/06/10 by Dauphin, Yann, Pascanu, Razvan, Gulcehre, Caglar +3 · 34 citations
#FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Optimization and Control (math.OC)
- Progress & Compress: A scalable framework for continual learning
2018/05/16 by Schwarz, Jonathan, Luketina, Jelena, Czarnecki, Wojciech M. +4 · 37 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Natural Neural Networks
2015/07/01 by Guillaume Desjardins, Desjardins, Guillaume, Karen Simonyan +5 · 1 voice · 14 citations
Computer Science · #Neural Networks and Applications #cs.LG #cs.NE #stat.ML
- Round and Round We Go! What makes Rotary Positional Encodings useful?
2024/10/08 by Federico Barbero, Barbero, Federico, Alex Vitvitskyi +7 · 4 voices · 24 citations
Computer Science · #Constraint Satisfaction and Optimization #Natural Language Processing Techniques #Speech and dialogue systems #cs.CL #cs.LG
- Theano: new features and speed improvements
2012/11/23 by Frédéric Bastien, Pascal Lamblin, Bastien, Frédéric +15 · 47 citations
Computer Science · Physics and Astronomy · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Numerical Methods and Algorithms #Parallel Computing and Optimization Techniques #Symbolic Computation (cs.SC)
- Learning to Navigate in Complex Environments
2016/11/11 by Piotr Mirowski, Mirowski, Piotr, Razvan Pascanu +21 · 36 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Robotic Path Planning Algorithms #Robotics (cs.RO)
- Model compression via distillation and quantization
2018/02/15 by Antonio Polino, Polino, Antonio, Razvan Pascanu +3 · 39 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
- Interaction Networks for Learning about Objects, Relations and Physics
2016/12/01 by Battaglia, Peter W., Pascanu, Razvan, Lai, Matthew +2 · 27 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Why do LLMs attend to the first token?
2025/04/03 by Federico Barbero, Álvaro Arroyo, Barbero, Federico +11 · 3 voices · 34 citations
#cs.CL
- Stabilizing Transformers for Reinforcement Learning
2019/10/13 by Emilio Parisotto, Hao Song, Parisotto, Emilio +23 · 33 citations
Computer Science · Neuroscience · #Reinforcement Learning in Robotics #Evolutionary Algorithms and Applications #Neural dynamics and brain function
- On the number of response regions of deep feed forward networks with\n piece-wise linear activations
2013/12/20 by Razvan Pascanu, Guido Montúfar, Pascanu, Razvan +3 · 26 citations
Engineering · Neuroscience · #Advanced Memory and Neural Computing #Neural dynamics and brain function #Ferroelectric and Negative Capacitance Devices
- Revisiting Natural Gradient for Deep Networks
2013/01/16 by Razvan Pascanu, Pascanu, Razvan, Yoshua Bengio +1 · 20 citations
Computer Science · #Stochastic Gradient Optimization Techniques
- Understanding plasticity in neural networks
2023/03/02 by Clare Lyle, Zeyu Zheng, Lyle, Clare +9 · 29 citations
Computer Science · Engineering · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
- Meta-Learning with Latent Embedding Optimization
2018/07/16 by Andrei A. Rusu, Dushyant Rao, Rusu, Andrei A. +11 · 24 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Hyperbolic Attention Networks
2018/05/24 by Çağlar Gülçehre, Gulcehre, Caglar, Misha Denil +19 · 18 citations
Computer Science · #Advanced Graph Neural Networks #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Neural and Evolutionary Computing (cs.NE) #Topic Modeling
- Distral: Robust Multitask Reinforcement Learning
2017/07/13 by Yee Whye Teh, Teh, Yee Whye, Victor Bapst +13 · 16 citations
Computer Science · #Adaptive Dynamic Programming Control #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM #Reinforcement Learning in Robotics
- Imagination-Augmented Agents for Deep Reinforcement Learning
2017/07/19 by Théophane Weber, Weber, Théophane, Sébastien Racanière +26 · 16 citations
Computer Science · #Reinforcement Learning in Robotics
- Functional Regularisation for Continual Learning with Gaussian Processes
2019/01/31 by Michalis K. Titsias, Jonathan Schwarz, Titsias, Michalis K. +7 · 16 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms
- Understanding the Role of Training Regimes in Continual Learning
2020/06/12 by Mirzadeh, Seyed Iman, Farajtabar, Mehrdad, Pascanu, Razvan +1 · 12 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- Linear Mode Connectivity in Multitask and Continual Learning
2020/10/09 by Seyed Iman Mirzadeh, Mehrdad Farajtabar, Mirzadeh, Seyed Iman +7 · 12 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Sparse and Compressive Sensing Techniques
- On the saddle point problem for non-convex optimization
2014/05/19 by Razvan Pascanu, Yann Dauphin, Pascanu, Razvan +5 · 13 citations
Computer Science · Engineering · #Stochastic Gradient Optimization Techniques #Sparse and Compressive Sensing Techniques #Gaussian Processes and Bayesian Inference
- Visual Interaction Networks
2017/06/05 by Nicholas Watters, Andrea Tacchetti, Watters, Nicholas +9 · 17 citations
Computer Science · #Data Visualization and Analytics #Human Pose and Action Recognition #Anomaly Detection Techniques and Applications
- Softmax is not Enough (for Sharp Size Generalisation)
2024/10/01 by Petar Veličković, Veličković, Petar, Christos Perivolaropoulos +5 · 6 voices · 8 citations
Decision Sciences · #Scientific Computing and Data Management #cs.AI #cs.IT #cs.LG
- BYOL works even without batch statistics
2020/10/20 by Pierre H. Richemond, Richemond, Pierre H., Jean-Bastien Grill +19 · 1 voice · 3 citations
Computer Science · Mathematics · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #cs.CV #cs.LG #stat.ML
- A study on the plasticity of neural networks
2021/05/31 by Tudor Berariu, Wojciech Marian Czarnecki, Berariu, Tudor +11 · 11 citations
Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Model Reduction and Neural Networks
- Transformers need glasses! Information over-squashing in language tasks
2024/06/06 by Federico Barbero, Andrea Banino, Barbero, Federico +13 · 18 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech and dialogue systems
- Distilling Policy Distillation
2019/02/06 by Wojciech Marian Czarnecki, Czarnecki, Wojciech Marian, Razvan Pascanu +9 · 9 citations
Computer Science · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Reinforcement Learning in Robotics
- Meta-learning of Sequential Strategies
2019/05/08 by Pedro A. Ortega, Jane X. Wang, Ortega, Pedro A. +45 · 13 citations
Computer Science · #Machine Learning and Data Classification #Gaussian Processes and Bayesian Inference #Data Stream Mining Techniques
- Continual Unsupervised Representation Learning
2019/10/31 by Dushyant Rao, Francesco Visin, Rao, Dushyant +9 · 13 citations
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Healthcare
- Adapting Auxiliary Losses Using Gradient Similarity
2018/12/05 by Yunshu Du, Wojciech Marian Czarnecki, Du, Yunshu +8 · 10 citations
Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
- Disentangling the Causes of Plasticity Loss in Neural Networks
2024/02/29 by Lyle, Clare, Zheng, Zeyu, Khetarpal, Khimya +4 · 15 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Sim-to-Real Robot Learning from Pixels with Progressive Nets
2016/10/13 by Rusu, Andrei A., Vecerik, Mel, Rothörl, Thomas +3 · 6 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Robotics (cs.RO)
- Improving fine-grained understanding in image-text pre-training
2024/01/18 by Bica, Ioana, Ilić, Anastasija, Bauer, Matthias +8 · 12 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Pre-training via Denoising for Molecular Property Prediction
2022/05/31 by Zaidi, Sheheryar, Schaarschmidt, Michael, Martens, James +6 · 8 citations
#Biomolecules (q-bio.BM) #FOS: Biological sciences #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- On the generalization of language models from in-context learning and finetuning: a controlled study
2025/05/01 by Andrew K. Lampinen, Lampinen, Andrew K., Arslan Chaudhry +18 · 4 voices · 12 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.LG
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
2025/04/22 by Thomas Schmied, Schmied, Thomas, Jörg Bornschein +7 · 5 voices · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.LG
- Normalization and effective learning rates in reinforcement learning
2024/07/01 by Clare Lyle, Zeyu Zheng, Lyle, Clare +11 · 12 citations
Computer Science · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Learned-Norm Pooling for Deep Feedforward and Recurrent Neural Networks
2013/11/07 by Çağlar Gülçehre, Gulcehre, Caglar, Kyunghyun Cho +5 · 4 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
- Pushing the limits of self-supervised ResNets: Can we outperform supervised learning without labels on ImageNet?
2022/01/13 by Nenad Tomašev, Tomasev, Nenad, Ioana Bica +11 · 6 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications
- The CLRS Algorithmic Reasoning Benchmark
2022/05/31 by Petar Veličković, Veličković, Petar, Adrià Puigdomènech Badia +13 · 6 citations
Computer Science · #Fuzzy Logic and Control Systems
- Relational Deep Reinforcement Learning
2018/06/05 by Vinícius Zambaldi, Zambaldi, Vinicius, David Raposo +29 · 5 citations
Computer Science · #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics
- Continual World: A Robotic Benchmark For Continual Reinforcement\n Learning
2021/05/23 by Maciej Wołczyk, Michał Zając, Wołczyk, Maciej +7 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Robotics (cs.RO)
- Wide Neural Networks Forget Less Catastrophically
2021/10/21 by Seyed Iman Mirzadeh, Arslan Chaudhry, Mirzadeh, Seyed Iman +10 · 5 citations
Computer Science · Medicine · #Artificial Intelligence (cs.AI) #COVID-19 diagnosis using AI #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
2024/05/01 by Skander Moalla, A. Miele, Moalla, Skander +6 · 8 citations
Business, Management and Accounting · #Outsourcing and Supply Chain Management
- Agency Is Frame-Dependent
2025/02/06 by David Abel, Abel, David, André Barreto +30 · 4 voices · 1 citation
Neuroscience · Psychology · #Action Observation and Synchronization #Embodied and Extended Cognition #Free Will and Agency #cs.AI
- Behavior Priors for Efficient Reinforcement Learning
2020/10/27 by Dhruva Tirumala, Alexandre Galashov, Tirumala, Dhruva +19 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Mobile Crowdsensing and Crowdsourcing #Reinforcement Learning in Robotics
- Learning Deep Generative Models of Graphs
2018/03/08 by Li, Yujia, Vinyals, Oriol, Dyer, Chris +2 · 3 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
2024/10/29 by Thomas Schmied, Schmied, Thomas, Thomas Adler +16 · 2 voices · 2 citations
Computer Science · #Adversarial Robustness in Machine Learning #Explainable Artificial Intelligence (XAI) #cs.AI #cs.LG
- Advances in Optimizing Recurrent Networks
2012/12/04 by Yoshua Bengio, Bengio, Yoshua, Nicolas Boulanger-Lewandowski +3 · 3 citations
Computer Science · #Music and Audio Processing #Neural Networks and Applications #Computational Physics and Python Applications
- Discovering modular solutions that generalize compositionally
2023/12/22 by Simon Schug, Seijin Kobayashi, Schug, Simon +15 · 6 citations
Computer Science · #Topic Modeling #Explainable Artificial Intelligence (XAI)
- Learning to Modulate pre-trained Models in RL
2023/06/26 by Thomas Schmied, Schmied, Thomas, Markus Hofmarcher +7 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics
- Deep Reinforcement Learning with Plasticity Injection
2023/05/24 by Nikishin, Evgenii, Oh, Junhyuk, Ostrovski, Georg +4 · 5 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- The Tunnel Effect: Building Data Representations in Deep Neural Networks
2023/05/31 by Wojciech Masarczyk, Mateusz Ostaszewski, Masarczyk, Wojciech +9 · 4 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Spectral Normalisation for Deep Reinforcement Learning: an Optimisation Perspective
2021/05/11 by Florin Gogianu, Gogianu, Florin, Tudor Berariu +9 · 3 citations
Computer Science · #Reinforcement Learning in Robotics #Adaptive Dynamic Programming Control #Machine Learning and ELM
- Theano: A Python framework for fast computation of mathematical expressions
2016/05/09 by The Theano Development Team, Al-Rfou, Rami, Alain, Guillaume +110 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Mathematical Software (cs.MS) #Symbolic Computation (cs.SC)
- How do language models learn facts? Dynamics, curricula and hallucinations
2025/03/27 by Nicolas Zucchet, Jörg Bornschein, Zucchet, Nicolas +9 · 2 voices · 8 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques
- Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation Problem
2024/02/05 by Maciej Wołczyk, Bartłomiej Cupiał, Wołczyk, Maciej +13 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics
- Disentangling Transfer in Continual Reinforcement Learning
2022/09/28 by Maciej Wołczyk, Michał Zając, Wołczyk, Maciej +7 · 3 citations
Computer Science · #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Mix&Match - Agent Curricula for Reinforcement Learning
2018/06/05 by Wojciech Marian Czarnecki, Siddhant M. Jayakumar, Czarnecki, Wojciech Marian +13 · 2 citations
Computer Science · #Reinforcement Learning in Robotics #Multi-Agent Systems and Negotiation #Software Engineering Research
- MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
2025/06/05 by Johannes von Oswald, von Oswald, Johannes, Nino Scherrer +30 · 11 citations
Computer Science · #Topic Modeling #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning
- Meta-Learning with Warped Gradient Descent
2019/08/30 by Flennerhag, Sebastian, Rusu, Andrei A., Pascanu, Razvan +3 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
2023/07/21 by Antonio Orvieto, Orvieto, Antonio, Soham De +7 · 4 citations
Computer Science · #Algorithms and Data Compression #Error Correcting Code Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications #Neural and Evolutionary Computing (cs.NE)
- Optimizers Qualitatively Alter Solutions And We Should Leverage This
2025/07/16 by Razvan Pascanu, Clare Lyle, Pascanu, Razvan +17 · 2 voices · 7 citations
Computer Science · #Stochastic Gradient Optimization Techniques #Advanced Neural Network Applications #Explainable Artificial Intelligence (XAI)
- Regularized Behavior Value Estimation
2021/03/17 by Gulcehre, Caglar, Colmenarejo, Sergio Gómez, Wang, Ziyu +7 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Attention as a Hypernetwork
2024/06/09 by Schug, Simon, Kobayashi, Seijin, Akram, Yassir +2 · 4 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- Information asymmetry in KL-regularized RL
2019/05/03 by Alexandre Galashov, Siddhant M. Jayakumar, Galashov, Alexandre +17 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Neural Networks and Applications #Reinforcement Learning in Robotics
- Temporal Difference Uncertainties as a Signal for Exploration
2020/10/05 by Sebastian Flennerhag, Jane X. Wang, Flennerhag, Sebastian +17 · 3 citations
Computer Science · Decision Sciences · #Advanced Multi-Objective Optimization Algorithms #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Simulation Techniques and Applications
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
2019/03/18 by Dhruva Tirumala, Hyeonwoo Noh, Tirumala, Dhruva +15 · 4 citations
Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Advanced Bandit Algorithms Research #Machine Learning and Algorithms
- When Does Re-initialization Work?
2022/06/20 by Zaidi, Sheheryar, Berariu, Tudor, Kim, Hyunjik +4 · 2 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- An Empirical Study of Implicit Regularization in Deep Offline RL
2022/07/05 by Çağlar Gülçehre, Srivatsan Srinivasan, Gulcehre, Caglar +13 · 2 citations
Computer Science · #Reinforcement Learning in Robotics #Machine Learning and ELM #Stochastic Gradient Optimization Techniques
- Plasticity as the Mirror of Empowerment
2025/05/15 by David Abel, Abel, David, Michael Bowling +30 · 3 voices · 3 citations
Neuroscience · Psychology · #cs.AI #cs.LG
- How to Construct Deep Recurrent Neural Networks
2013/12/20 by Pascanu, Razvan, Gulcehre, Caglar, Cho, Kyunghyun +1 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- Deep Grokking: Would Deep Neural Networks Generalize Better?
2024/05/29 by Fan, Simin, Pascanu, Razvan, Jaggi, Martin · 3 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Retrieval-Augmented Decision Transformer: External Memory for In-context RL
2024/10/09 by Schmied, Thomas, Paischer, Fabian, Patil, Vihang +3 · 4 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Asynchronous Algorithmic Alignment with Cocycles
2023/06/27 by Dudzik, Andrew, von Glehn, Tamara, Pascanu, Razvan +1 · 2 citations
#Artificial Intelligence (cs.AI) #Commutative Algebra (math.AC) #Data Structures and Algorithms (cs.DS) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG)
- Local minima in training of neural networks
2016/11/19 by Swirszcz, Grzegorz, Czarnecki, Wojciech Marian, Pascanu, Razvan · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- Been There, Done That: Meta-Learning with Episodic Recall
2018/05/24 by Ritter, Samuel, Wang, Jane X., Kurth-Nelson, Zeb +4 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- TRecViT: A Recurrent Video Transformer
2024/12/18 by Viorica Pătrăucean, Pătrăucean, Viorica, Xu Owen He +23 · 1 voice · 2 citations
Computer Science · #Advanced Vision and Imaging #Video Analysis and Summarization #Video Coding and Compression Technologies
- Relational recurrent neural networks
2018/06/05 by Santoro, Adam, Faulkner, Ryan, Raposo, David +7 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Deep Self-Taught Learning for Handwritten Character Recognition
2010/09/18 by Frédéric Bastien, Bastien, Frédéric, Yoshua Bengio +31 · 1 citation
Computer Science · Engineering · #68T05 #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #I.2.6 #Image Processing and 3D Reconstruction #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE) #Vehicle License Plate Recognition
- Ray Interference: a Source of Plateaus in Deep Reinforcement Learning
2019/04/25 by Schaul, Tom, Borsa, Diana, Modayil, Joseph +1 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Task Agnostic Continual Learning via Meta Learning
2019/06/12 by He, Xu, Sygnowski, Jakub, Galashov, Alexandre +3 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE)
- From Markov to Laplace: How Mamba In-Context Learns Markov Chains
2025/02/14 by Bondaschi, Marco, Rajaraman, Nived, Wei, Xiuying +5 · 4 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG)
- Block Mean Approximation for Efficient Second Order Optimization
2018/04/16 by Lu, Yao, Harandi, Mehrtash, Hartley, Richard +1 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Drawing Multiple Augmentation Samples Per Image During Training\n Efficiently Decreases Test Error
2021/05/27 by Stanislav Fort, Andrew Brock, Fort, Stanislav +7 · 2 citations
Computer Science · Engineering · #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications #Medical Imaging and Analysis
- Transformers meet Neural Algorithmic Reasoners
2024/06/13 by Wilfried Bounsi, Borja Ibarz, Bounsi, Wilfried +13 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
- Task-agnostic Continual Learning with Hybrid Probabilistic Models
2021/06/24 by Polina Kirichenko, Mehrdad Farajtabar, Kirichenko, Polina +15 · 1 citation
Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and ELM
- Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
2024/07/13 by Xiuying Wei, Skander Moalla, Wei, Xiuying +5 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Architecture Matters in Continual Learning
2022/02/01 by Mirzadeh, Seyed Iman, Chaudhry, Arslan, Yin, Dong +4 · 1 citation
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Towards Robust and Efficient Continual Language Learning
2023/07/11 by Adam Fisch, Amal Rannen-Triki, Fisch, Adam +11 · 1 citation
Computer Science · #Domain Adaptation and Few-Shot Learning #Topic Modeling
- Hadamard product in deep learning: Introduction, Advances and Challenges
2025/04/17 by Grigorios G. Chrysos, Chrysos, Grigorios G, Yongtao Wu +7 · 3 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Neural Network Applications #Mobile Crowdsensing and Crowdsourcing
- Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
2024/03/12 by Simon Dufort-Labbé, Pierluca D’Oro, Dufort-Labbé, Simon +9 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural Networks and Applications
- What Can Grokking Teach Us About Learning Under Nonstationarity?
2025/07/26 by Lyle, Clare, Sokar, Gharda, Pascanu, Razvan +1 · 2 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
- When can transformers compositionally generalize in-context?
2024/07/17 by Kobayashi, Seijin, Schug, Simon, Akram, Yassir +5 · 1 citation
#FOS: Computer and information sciences #Machine Learning (cs.LG) #Neural and Evolutionary Computing (cs.NE)
- Reasoning-Modulated Representations
2021/07/19 by Petar Veličković, Veličković, Petar, Matko Bošnjak +11 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Topic Modeling
- How Overconfidence in Initial Choices and Underconfidence Under Criticism Modulate Change of Mind in Large Language Models
2025/07/03 by Kumaran, Dharshan, Fleming, Stephen M, Markeeva, Larisa +8 · 2 citations
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)