2025/12/05 by Jiahao You, You, Jiahao, Ziye Jia +9
Computer Science · Engineering · #Adaptability #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Base station #Convergence (economics) #FOS: Computer and information sciences #IoT and Edge/Fog Computing #Key (lock) #Machine Learning (cs.LG) #Markov decision process #Optimization problem #Reinforcement learning #Task (project management) #Trajectory #Trajectory optimization #UAV Applications and Optimization
paper · pdf · doi:10.48550/arxiv.2512.11862
published in arXiv (Cornell University) (Cornell University)
openalex publication_date 2025/12/05 · openalex created_date 2025/12/17 · openalex updated_date 2026/08/05
The low-altitude intelligent networks (LAINs) emerge as a promising architecture for delivering low-latency and energy-efficient edge intelligence in dynamic and infrastructure-limited environments. By integrating unmanned aerial vehicles (UAVs), aerial base stations, and terrestrial base stations, LAINs can support mission-critical applications such as disaster response, environmental monitoring, and real-time sensing. However, these systems face key challenges, including energy-constrained UAVs, stochastic task arrivals, and heterogeneous computing resources. To address these issues, we propose an integrated air-ground collaborative network and formulate a time-dependent integer nonlinear programming problem that jointly optimizes UAV trajectory planning and task offloading decisions. The problem is challenging to solve due to temporal coupling among decision variables. Therefore, we design a hierarchical learning framework with two timescales. At the large timescale, a Vickrey-Clarke-Groves auction mechanism enables the energy-aware and incentive-compatible trajectory assignment. At the small timescale, we propose the diffusion-heterogeneous-agent proximal policy optimization, a generative multi-agent reinforcement learning algorithm that embeds latent diffusion models into actor networks. Each UAV samples actions from a Gaussian prior and refines them via observation-conditioned denoising, enhancing adaptability and policy diversity. Extensive simulations show that our framework outperforms baselines in energy efficiency, task success rate, and convergence performance.