2020/06/29 by Zahi M. Kakish, Zahi Kakish, Karthik Elamvazhuthi +4 · 2 citations
Computer Science · Engineering · Mathematics · Physics and Astronomy · Social Sciences · #Artificial Intelligence (cs.AI) #Artificial intelligence #Computer science #Control (management) #Distribution (mathematics) #FOS: Computer and information sciences #Graph #Grid #Machine Learning (cs.LG) #Mathematical optimization #Mathematics #Multi-agent system #Multiagent Systems (cs.MA) #Opinion Dynamics and Social Influence #Reinforcement learning #Robot #Robotics (cs.RO) #Swarm behaviour #Swarm robotics #Theoretical computer science #Transportation Planning and Optimization #Transportation and Mobility Innovations #Vertex (graph theory) #cs.AI #cs.LG #cs.MA #cs.RO
paper · pdf · doi:10.48550/arxiv.2006.15807
published in arXiv (Cornell University) (Cornell University) · Paper was submitted to Conference on Robot Learning 2019 and IEEE Robotics and Automation Letters 2020 Revised, updated, and submitted to DARS/SWARMS 2021
openalex publication_date 2020/06/29 · arxiv created 2020/12/12 · arxiv updated 2020/12/15 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
In this paper, we present a reinforcement learning approach to designing a\ncontrol policy for a "leader" agent that herds a swarm of "follower" agents,\nvia repulsive interactions, as quickly as possible to a target probability\ndistribution over a strongly connected graph. The leader control policy is a\nfunction of the swarm distribution, which evolves over time according to a\nmean-field model in the form of an ordinary difference equation. The dependence\nof the policy on agent populations at each graph vertex, rather than on\nindividual agent activity, simplifies the observations required by the leader\nand enables the control strategy to scale with the number of agents. Two\nTemporal-Difference learning algorithms, SARSA and Q-Learning, are used to\ngenerate the leader control policy based on the follower agent distribution and\nthe leader's location on the graph. A simulation environment corresponding to a\ngrid graph with 4 vertices was used to train and validate the control policies\nfor follower agent populations ranging from 10 to 100. Finally, the control\npolicies trained on 100 simulated agents were used to successfully redistribute\na physical swarm of 10 small robots to a target distribution among 4 spatial\nregions.\n