2021/03/14 by Mohammad G. Khoshkholgh, Halim Yanıkömeroğlu, Khoshkholgh, Mohammad Ghadir +1
Engineering · #Air Traffic Management and Optimization #FOS: Computer and information sciences #Indoor and Outdoor Localization Technologies #Information Theory (cs.IT) #UAV Applications and Optimization
paper · pdf · doi:10.48550/arxiv.2103.08034
openalex publication_date 2021/03/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We address the mobility management of an autonomous UAV-mounted base station\n(UAV-BS) that provides communication services to a cluster of users on the\nground while the geographical characteristics (e.g., location and boundary) of\nthe cluster, the geographical locations of the users, and the characteristics\nof the radio environment are unknown. UAVBS solely exploits the received signal\nstrengths (RSS) from the users and accordingly chooses its (continuous) 3-D\nspeed to constructively navigate, i.e., improving the transmitted data rate. To\ncompensate for the lack of a model, we adopt policy gradient deep reinforcement\nlearning. As our approach does not rely on any particular information about the\nusers as well as the radio environment, it is flexible and respects the privacy\nconcerns. Our experiments indicate that despite the minimum available\ninformation the UAV-BS is able to distinguish between high-rise (often\nnon-line-of-sight dominant) and sub-urban (mainly line-of-sight dominant)\nenvironments such that in the former (resp. latter) it tends to reduce (resp.\nincrease) its height and stays close (resp. far) to the cluster. We further\nobserve that the choice of the reward function affects the speed and the\nability of the agent to adhere to the problem constraints without affecting the\ndelivered data rate.\n