2025/05/06 by Yue Chen, Chen, Yue, Hui Kang +15 · 3 citations
Computer Science · Engineering · #Battery (electricity) #Computational complexity theory #Convergence (economics) #Distributed Control Multi-Agent Systems #Energy management #Infrared Target Detection Methodologies #Markov decision process #Process (computing) #Reinforcement learning #Resource allocation #Resource management (computing) #UAV Applications and Optimization #Wireless
paper · pdf · doi:10.48550/arxiv.2505.03230
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/05/06 · openalex created_date 2025/10/16 · openalex updated_date 2026/08/05
The integration of simultaneous wireless information and power transfer (SWIPT) technology in 6G Internet of Things (IoT) networks faces significant challenges in remote areas and disaster scenarios where ground infrastructure is unavailable. This paper proposes a novel unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) system enhanced by directional antennas to provide both computational resources and energy support for ground IoT terminals. However, such systems require multiple trade-off policies to balance UAV energy consumption, terminal battery levels, and computational resource allocation under various constraints, including limited UAV battery capacity, non-linear energy harvesting characteristics, and dynamic task arrivals. To address these challenges comprehensively, we formulate a bi-objective optimization problem that simultaneously considers system energy efficiency and terminal battery sustainability. We then reformulate this non-convex problem with a hybrid solution space as a Markov decision process (MDP) and propose an improved soft actor-critic (SAC) algorithm with an action simplification mechanism to enhance its convergence and generalization capabilities. Simulation results have demonstrated that our proposed approach outperforms various baselines in different scenarios, achieving efficient energy management while maintaining high computational performance. Furthermore, our method shows strong generalization ability across different scenarios, particularly in complex environments, validating the effectiveness of our designed boundary penalty and charging reward mechanisms.