2021/07/29 by Nicholas Kullman, Nicholas D. Kullman, Martin Cousineau +2
Business, Management and Accounting · Engineering · #Electric Vehicles and Infrastructure #Sharing Economy and Platforms #Transportation and Mobility Innovations
paper · doi:10.1287/trsc.2021.1042
openalex publication_date 2021/07/29 · crossref created 2021/07/29 · crossref issued 2022/05/01 · crossref published 2022/05/01 · crossref published-print 2022/05/01 · crossref deposited 2023/04/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/22 · crossref indexed 2026/08/01
We consider the problem of an operator controlling a fleet of electric vehicles for use in a ride-hailing service. The operator, seeking to maximize profit, must assign vehicles to requests as they arise as well as recharge and reposition vehicles in anticipation of future requests. To solve this problem, we employ deep reinforcement learning, developing policies whose decision making uses [Formula: see text]-value approximations learned by deep neural networks. We compare these policies against a reoptimization-based policy and against dual bounds on the value of an optimal policy, including the value of an optimal policy with perfect information, which we establish using a Benders-based decomposition. We assess performance on instances derived from real data for the island of Manhattan in New York City. We find that, across instances of varying size, our best policy trained with deep reinforcement learning outperforms the reoptimization approach. We also provide evidence that this policy may be effectively scaled and deployed on larger instances without retraining.