vix.ing · top · new · best · stats

Worst-Case Control and Learning Using Partial Observations Over an Infinite Time-Horizon

2023/03/28 by Aditya Dave, Dave, Aditya, Ioannis Faros +5
Computer Science · Decision Sciences · Engineering · Mathematics · #Advanced Statistical Process Monitoring #Algorithm #Artificial Intelligence (cs.AI) #Artificial intelligence #Class (philosophy) #Computer science #Control (management) #Dynamic programming #FOS: Computer and information sciences #FOS: Electrical engineering #FOS: Mathematics #Fault Detection and Control Systems #Mathematical optimization #Mathematics #Observable #Optimal control #Optimization and Control (math.OC) #Software Reliability and Analysis Research #State (computer science) #Systems and Control (eess.SY) #Time horizon #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2303.16321

published in arXiv (Cornell University) (Cornell University)

openalex publication_date 2023/03/28 · openalex created_date 2023/04/05 · openalex updated_date 2026/07/28

Abstract

Safety-critical cyber-physical systems require control strategies whose worst-case performance is robust against adversarial disturbances and modeling uncertainties. In this paper, we present a framework for approximate control and learning in partially observed systems to minimize the worst-case discounted cost over an infinite time horizon. We model disturbances to the system as finite-valued uncertain variables with unknown probability distributions. For problems with known system dynamics, we construct a dynamic programming (DP) decomposition to compute the optimal control strategy. Our first contribution is to define information states that improve the computational tractability of this DP without loss of optimality. Then, we describe a simplification for a class of problems where the incurred cost is observable at each time instance. Our second contribution is defining an approximate information state that can be constructed or learned directly from observed data for problems with observable costs. We derive bounds on the performance loss of the resulting approximate control strategy and illustrate the effectiveness of our approach in partially observed decision-making problems with a numerical example.

Related