2022/02/14 by Jiawei Huang, Huang, Jiawei, Jing‐Lin Chen +9 · 1 citation
Computer Science · Engineering · #Age of Information Optimization #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Modular Robots and Swarm Intelligence #Optimization and Search Problems
paper · pdf · doi:10.48550/arxiv.2202.06450
openalex publication_date 2022/02/14 · openalex created_date 2022/04/03 · openalex updated_date 2026/07/28
Deployment efficiency is an important criterion for many real-world applications of reinforcement learning (RL). Despite the community's increasing interest, there lacks a formal theoretical formulation for the problem. In this paper, we propose such a formulation for deployment-efficient RL (DE-RL) from an "optimization with constraints" perspective: we are interested in exploring an MDP and obtaining a near-optimal policy within minimal deployment complexity, whereas in each deployment the policy can sample a large batch of data. Using finite-horizon linear MDPs as a concrete structural model, we reveal the fundamental limit in achieving deployment efficiency by establishing information-theoretic lower bounds, and provide algorithms that achieve the optimal deployment efficiency. Moreover, our formulation for DE-RL is flexible and can serve as a building block for other practically relevant settings; we give "Safe DE-RL" and "Sample-Efficient DE-RL" as two examples, which may be worth future investigation.