vix.ing · top · new · best · stats · spec

Zanette, Andrea

  1. ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
    2024/02/29 by Yifei Zhou, Andrea Zanette, Zhou, Yifei +7 · 37 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  2. Training Language Models to Reason Efficiently
    2025/02/06 by Daman Arora, Arora, Daman, Andrea Zanette +1 · 65 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  3. Fast Best-of-N Decoding via Speculative Rejection
    2024/10/26 by Hanshi Sun, Sun, Hanshi, Momin Haider +13 · 38 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cellular Automata and Applications #Computation and Language (cs.CL) #DNA and Biological Computing #Error Correcting Code Techniques #FOS: Computer and information sciences
  4. Learning Near Optimal Policies with Low Inherent Bellman Error
    2020/02/29 by Andrea Zanette, Alessandro Lazaric, Zanette, Andrea +5 · 13 citations
    Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Smart Grid Energy Management
  5. Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
    2021/08/19 by Andrea Zanette, Martin J. Wainwright, Zanette, Andrea +3 · 6 citations
    Computer Science · #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics
  6. Frequentist Regret Bounds for Randomized Least-Squares Value Iteration
    2019/11/01 by Zanette, Andrea, Brandfonbrener, David, Brunskill, Emma +2 · 3 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  7. Exponential Lower Bounds for Batch Reinforcement Learning: Batch RL can be Exponentially Harder than Online RL
    2020/12/14 by Zanette, Andrea · 3 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  8. Can Large Reasoning Models Self-Train?
    2025/05/27 by Shafayat, Sheikh, Tajwar, Fahim, Salakhutdinov, Ruslan +2 · 15 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. SPEED-RL: Faster Training of Reasoning Models via Online Curriculum Learning
    2025/06/10 by Zhang, Ruiqi, Arora, Daman, Mei, Song +1 · 13 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG)
  10. Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline Data
    2023/07/10 by Zhang, Ruiqi, Zanette, Andrea · 2 citations
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
  11. When is Realizability Sufficient for Off-Policy Reinforcement Learning?
    2022/11/10 by Andrea Zanette, Zanette, Andrea · 1 citation
    Computer Science · Decision Sciences · Biochemistry, Genetics and Molecular Biology · #Reinforcement Learning in Robotics #Advanced Bandit Algorithms Research #Gene Regulatory Network Analysis
  12. Accelerating Unbiased LLM Evaluation via Synthetic Feedback
    2025/02/14 by Zhaoyi Zhou, Yuda Song, Zhou, Zhaoyi +3 · 1 citation
    Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Non-Destructive Testing Techniques