Bilal Piot
- Mastering the game of Stratego with model-free multiagent reinforcement learning
2022/06/30 by Julien Perolat, Julien Pérolat, Bart De Vylder +38 · 6 voices · 21 citations
Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Artificial Intelligence in Games #Advanced Bandit Algorithms Research
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
2025/07/07 by Gheorghe Comanici, Comanici, Gheorghe, Eric Bieber +6844 · 8 voices · 1347 citations
#cs.CL #cs.AI
- Noisy Networks for Exploration
2017/06/30 by Meire Fortunato, Fortunato, Meire, Mohammad Gheshlaghi Azar +22 · 2 voices · 40 citations
Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Adversarial Robustness in Machine Learning #Advanced Bandit Algorithms Research
- Bootstrap your own latent: A new approach to self-supervised Learning
2020/06/13 by Jean-Bastien Grill, Grill, Jean-Bastien, Florian Strub +25 · 288 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification
- A General Theoretical Paradigm to Understand Learning from Human Preferences
2023/10/18 by Mohammad Gheshlaghi Azar, Mark Rowland, Azar, Mohammad Gheshlaghi +11 · 3 voices · 116 citations
Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Advanced Bandit Algorithms Research
- Direct Language Model Alignment from Online AI Feedback
2024/02/07 by Shangmin Guo, Guo, Shangmin, Biao Zhang +23 · 1 voice · 36 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- Gemma 3 Technical Report
2025/03/25 by Gemma Team, Aishwarya Kamath, Johan Ferret +418 · 3 voices · 461 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
- Deep Q-learning from Demonstrations
2017/04/12 by Todd Hester, Matej Vecerik, Hester, Todd +27 · 2 voices · 13 citations
Computer Science · Economics, Econometrics and Finance · #Reinforcement Learning in Robotics #Software Engineering Research #Sports Analytics and Performance #cs.AI #cs.LG
- Gemma 2: Improving Open Language Models at a Practical Size
2024/07/31 by Morgane Rivière, Gemma Team, Riviere, Morgane +292 · 302 citations
Computer Science · #Natural Language Processing Techniques
- Rainbow: Combining Improvements in Deep Reinforcement Learning
2017/10/06 by Matteo Hessel, Hessel, Matteo, Joseph Modayil +17 · 88 citations
Computer Science · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
- Agent57: Outperforming the Atari Human Benchmark
2020/03/30 by Adrià Puigdomènech Badia, Bilal Piot, Badia, Adrià Puigdomènech +11 · 1 voice · 12 citations
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
- Shaking the foundations: delusions in sequence models for interaction and control
2021/10/20 by Pedro A. Ortega, Ortega, Pedro A., Markus Kunesch +37 · 2 voices · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.LG
- Gemma 4 Technical Report
2026/07/02 by Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320 · 7 voices · 8 citations
#cs.CL #cs.AI
- Nash Learning from Human Feedback
2023/12/01 by Rémi Munos, Michal Valko, Munos, Rémi +31 · 1 voice · 25 citations
#stat.ML #cs.AI #cs.GT #cs.LG #cs.MA
- Never Give Up: Learning Directed Exploration Strategies
2020/02/14 by Adrià Puigdomènech Badia, Pablo Sprechmann, Badia, Adrià Puigdomènech +18 · 12 citations
Computer Science · #Artificial Intelligence in Games #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
- Observational Learning by Reinforcement Learning
2017/06/20 by Diana Borsa, Borsa, Diana, Bilal Piot +5 · 1 voice · 1 citation
Computer Science · Engineering · #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Evolutionary Algorithms and Applications
- BYOL works even without batch statistics
2020/10/20 by Pierre H. Richemond, Richemond, Pierre H., Jean-Bastien Grill +19 · 1 voice · 3 citations
Computer Science · Mathematics · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.CV #cs.LG #stat.ML
- Multi-turn Reinforcement Learning from Preference Human Feedback
2024/05/23 by Lior Shani, Aviv Rosenberg, Shani, Lior +23 · 1 voice · 16 citations
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
- Generalized Preference Optimization: A Unified Approach to Offline Alignment
2024/02/08 by Yunhao Tang, Zhaohan Daniel Guo, Tang, Yunhao +17 · 12 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Constraint Satisfaction and Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multi-Criteria Decision Making
- Hindsight Credit Assignment
2019/12/05 by Anna Harutyunyan, Harutyunyan, Anna, Will Dabney +19 · 7 citations
Computer Science · Business, Management and Accounting · #Reinforcement Learning in Robotics #Financial Distress and Bankruptcy Prediction
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning
2020/04/30 by Daniel Guo, Bernardo Ávila Pires, Guo, Daniel +11 · 8 citations
Computer Science · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
- Neural Predictive Belief Representations
2018/11/15 by Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Guo, Zhaohan Daniel +7 · 5 citations
Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- Offline Regularised Reinforcement Learning for Large Language Models Alignment
2024/05/29 by Pierre Harvey Richemond, Richemond, Pierre Harvey, Yunhao Tang +33 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Learning from negative feedback, or positive feedback or both
2024/10/05 by Abbas Abdolmaleki, Bilal Piot, Abdolmaleki, Abbas +21 · 5 citations
Computer Science · #Semantic Web and Ontologies
- The Edge of Orthogonality: A Simple View of What Makes BYOL Tick
2023/02/09 by Pierre H. Richemond, Allison Tam, Richemond, Pierre H. +9 · 1 citation
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- Unlocking the Power of Representations in Long-term Novelty-based Exploration
2023/05/02 by Alaa Saade, Saade, Alaa, Steven Kapturowski +15 · 1 citation
Computer Science · #Anomaly Detection Techniques and Applications #Artificial Intelligence in Games #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting