Sam Work
- Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs
2025/03/18 by Nicolas Le Roux, Marc G. Bellemare, Roux, Nicolas Le +17 · 1 voice · 30 citations
Computer Science · Engineering · #Digital Rights Management and Security #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multi-Agent Systems and Negotiation #Scheduling and Optimization Algorithms #cs.LG