vix.ing · top · new · best · stats · spec

Definition and evaluation of model-free coordination of electrical\n vehicle charging with reinforcement learning

2018/09/27 by Nasrin Sadeghianpourhamami, Sadeghianpourhamami, Nasrin, Johannes Deleu +3
Engineering · #Advanced Battery Technologies Research #Artificial Intelligence (cs.AI) #Electric Vehicles and Infrastructure #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Transportation and Mobility Innovations

paper · pdf · doi:10.48550/arxiv.1809.10679

openalex publication_date 2018/09/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Initial DR studies mainly adopt model predictive control and thus require\naccurate models of the control problem (e.g., a customer behavior model), which\nare to a large extent uncertain for the EV scenario. Hence, model-free\napproaches, especially based on reinforcement learning (RL) are an attractive\nalternative. In this paper, we propose a new Markov decision process (MDP)\nformulation in the RL framework, to jointly coordinate a set of EV charging\nstations. State-of-the-art algorithms either focus on a single EV, or perform\nthe control of an aggregate of EVs in multiple steps (e.g., aggregate load\ndecisions in one step, then a step translating the aggregate decision to\nindividual connected EVs). On the contrary, we propose an RL approach to\njointly control the whole set of EVs at once. We contribute a new MDP\nformulation, with a scalable state representation that is independent of the\nnumber of EV charging stations. Further, we use a batch reinforcement learning\nalgorithm, i.e., an instance of fitted Q-iteration, to learn the optimal\ncharging policy. We analyze its performance using simulation experiments based\non a real-world EV charging data. More specifically, we (i) explore the various\nsettings in training the RL policy (e.g., duration of the period with training\ndata), (ii) compare its performance to an oracle all-knowing benchmark (which\nprovides an upper bound for performance, relying on information that is not\navailable or at least imperfect in practice), (iii) analyze performance over\ntime, over the course of a full year to evaluate possible performance\nfluctuations (e.g, across different seasons), and (iv) demonstrate the\ngeneralization capacity of a learned control policy to larger sets of charging\nstations.\n

Related