vix.ing · top · new · best · stats · spec

Bilal Piot

  1. Mastering the game of Stratego with model-free multiagent reinforcement learning
    2022/06/30 by Julien Perolat, Julien Pérolat, Bart De Vylder +38 · 6 voices · 21 citations
    Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Artificial Intelligence in Games #Advanced Bandit Algorithms Research
  2. Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
    2025/07/07 by Gheorghe Comanici, Comanici, Gheorghe, Eric Bieber +6844 · 8 voices · 1347 citations
    #cs.CL #cs.AI
  3. Noisy Networks for Exploration
    2017/06/30 by Meire Fortunato, Fortunato, Meire, Mohammad Gheshlaghi Azar +22 · 2 voices · 40 citations
    Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Adversarial Robustness in Machine Learning #Advanced Bandit Algorithms Research
  4. Bootstrap your own latent: A new approach to self-supervised Learning
    2020/06/13 by Jean-Bastien Grill, Grill, Jean-Bastien, Florian Strub +25 · 288 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification
  5. A General Theoretical Paradigm to Understand Learning from Human Preferences
    2023/10/18 by Mohammad Gheshlaghi Azar, Mark Rowland, Azar, Mohammad Gheshlaghi +11 · 3 voices · 116 citations
    Computer Science · Decision Sciences · #Reinforcement Learning in Robotics #Advanced Bandit Algorithms Research
  6. Direct Language Model Alignment from Online AI Feedback
    2024/02/07 by Shangmin Guo, Guo, Shangmin, Biao Zhang +23 · 1 voice · 36 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  7. Gemma 3 Technical Report
    2025/03/25 by Gemma Team, Aishwarya Kamath, Johan Ferret +418 · 3 voices · 461 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #cs.AI #cs.CL
  8. Deep Q-learning from Demonstrations
    2017/04/12 by Todd Hester, Matej Vecerik, Hester, Todd +27 · 2 voices · 13 citations
    Computer Science · Economics, Econometrics and Finance · #Reinforcement Learning in Robotics #Software Engineering Research #Sports Analytics and Performance #cs.AI #cs.LG
  9. Gemma 2: Improving Open Language Models at a Practical Size
    2024/07/31 by Morgane Rivière, Gemma Team, Riviere, Morgane +292 · 302 citations
    Computer Science · #Natural Language Processing Techniques
  10. Rainbow: Combining Improvements in Deep Reinforcement Learning
    2017/10/06 by Matteo Hessel, Hessel, Matteo, Joseph Modayil +17 · 88 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
  11. Agent57: Outperforming the Atari Human Benchmark
    2020/03/30 by Adrià Puigdomènech Badia, Bilal Piot, Badia, Adrià Puigdomènech +11 · 1 voice · 12 citations
    Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.LG #stat.ML
  12. Shaking the foundations: delusions in sequence models for interaction and control
    2021/10/20 by Pedro A. Ortega, Ortega, Pedro A., Markus Kunesch +37 · 2 voices · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Topic Modeling #cs.AI #cs.LG
  13. Gemma 4 Technical Report
    2026/07/02 by Gemma Team, Sherif El Abd, Vaibhav Aggarwal +320 · 7 voices · 8 citations
    #cs.CL #cs.AI
  14. Nash Learning from Human Feedback
    2023/12/01 by Rémi Munos, Michal Valko, Munos, Rémi +31 · 1 voice · 25 citations
    #stat.ML #cs.AI #cs.GT #cs.LG #cs.MA
  15. Never Give Up: Learning Directed Exploration Strategies
    2020/02/14 by Adrià Puigdomènech Badia, Pablo Sprechmann, Badia, Adrià Puigdomènech +18 · 12 citations
    Computer Science · #Artificial Intelligence in Games #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics
  16. Observational Learning by Reinforcement Learning
    2017/06/20 by Diana Borsa, Borsa, Diana, Bilal Piot +5 · 1 voice · 1 citation
    Computer Science · Engineering · #Reinforcement Learning in Robotics #Robot Manipulation and Learning #Evolutionary Algorithms and Applications
  17. BYOL works even without batch statistics
    2020/10/20 by Pierre H. Richemond, Richemond, Pierre H., Jean-Bastien Grill +19 · 1 voice · 3 citations
    Computer Science · Mathematics · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.CV #cs.LG #stat.ML
  18. Multi-turn Reinforcement Learning from Preference Human Feedback
    2024/05/23 by Lior Shani, Aviv Rosenberg, Shani, Lior +23 · 1 voice · 16 citations
    Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.LG
  19. Generalized Preference Optimization: A Unified Approach to Offline Alignment
    2024/02/08 by Yunhao Tang, Zhaohan Daniel Guo, Tang, Yunhao +17 · 12 citations
    Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Constraint Satisfaction and Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multi-Criteria Decision Making
  20. Hindsight Credit Assignment
    2019/12/05 by Anna Harutyunyan, Harutyunyan, Anna, Will Dabney +19 · 7 citations
    Computer Science · Business, Management and Accounting · #Reinforcement Learning in Robotics #Financial Distress and Bankruptcy Prediction
  21. Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning
    2020/04/30 by Daniel Guo, Bernardo Ávila Pires, Guo, Daniel +11 · 8 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Data Stream Mining Techniques #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
  22. Neural Predictive Belief Representations
    2018/11/15 by Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Guo, Zhaohan Daniel +7 · 5 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Domain Adaptation and Few-Shot Learning #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  23. Offline Regularised Reinforcement Learning for Large Language Models Alignment
    2024/05/29 by Pierre Harvey Richemond, Richemond, Pierre Harvey, Yunhao Tang +33 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  24. Learning from negative feedback, or positive feedback or both
    2024/10/05 by Abbas Abdolmaleki, Bilal Piot, Abdolmaleki, Abbas +21 · 5 citations
    Computer Science · #Semantic Web and Ontologies
  25. The Edge of Orthogonality: A Simple View of What Makes BYOL Tick
    2023/02/09 by Pierre H. Richemond, Allison Tam, Richemond, Pierre H. +9 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  26. Unlocking the Power of Representations in Long-term Novelty-based Exploration
    2023/05/02 by Alaa Saade, Saade, Alaa, Steven Kapturowski +15 · 1 citation
    Computer Science · #Anomaly Detection Techniques and Applications #Artificial Intelligence in Games #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Time Series Analysis and Forecasting