vix.ing · top · new · best · stats · spec

Policy optimization for CMDPs with bandit feedback: Best-of-both-worlds and beyond

2026/07/20 by Francesco Emanuele Stradi, Anna Lunghi, Matteo Castiglioni +2
Decision Sciences · Computer Science · #Advanced Bandit Algorithms Research #Reinforcement Learning in Robotics #Auction Theory and Applications

paper · doi:10.1016/j.artint.2026.104589

Citations