Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
2025/08/13 by Zhao, Changyuan, Liu, Guangyuan, Zhang, Ruichen +11 · 5 citations
#FOS: Computer and information sciences #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2508.09561
Abstract
Edge General Intelligence (EGI) represents a transformative evolution of edge computing, where distributed agents possess the capability to perceive, reason, and act autonomously across diverse, dynamic environments. Central to this vision are world models, which act as proactive internal simulators that not only predict but also actively imagine future trajectories, reason under uncertainty, and plan multi-step actions with foresight. This proactive nature allows agents to anticipate potential outcomes and optimize decisions ahead of real-world interactions. While prior works in robotics and gaming have showcased the potential of world models, their integration into the wireless edge for EGI remains underexplored. This survey bridges this gap by offering a comprehensive analysis of how world models can empower agentic artificial intelligence (AI) systems at the edge. We first examine the architectural foundations of world models, including latent representation learning, dynamics modeling, and imagination-based planning. Building on these core capabilities, we illustrate their proactive applications across EGI scenarios such as vehicular networks, unmanned aerial vehicle (UAV) networks, the Internet of Things (IoT) systems, and network functions virtualization, thereby highlighting how they can enhance optimization under latency, energy, and privacy constraints. We then explore their synergy with foundation models and digital twins, positioning world models as the cognitive backbone of EGI. Finally, we highlight open challenges, such as safety guarantees, efficient training, and constrained deployment, and outline future research directions. This survey provides both a conceptual foundation and a practical roadmap for realizing the next generation of intelligent, autonomous edge systems.
Citations
- Energy-Efficient RSMA-enabled Low-altitude MEC Optimization Via Generative AI-enhanced Deep Reinforcement Learning
- Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Simple, Good, Fast: Self-Supervised World Models Free of Baggage
- Sparse Imagination for Efficient Visual World Model Planning
- World Models for Cognitive Agents: Transforming Edge Intelligence in Future Networks
- AI Agents vs. Agentic AI: A Conceptual Taxonomy, Applications and Challenges
- A Comprehensive Survey of Large AI Models for Future Communications: Foundations, Applications and Challenges
- World Model-Based Learning for Long-Term Age of Information Minimization in Vehicular Networks
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions
- Generative AI-enabled Wireless Communications for Robust Low-Altitude Economy Networking
- Generative AI Enabled Robust Data Augmentation for Wireless Sensing in ISAC Networks
- A Survey of World Models for Autonomous Driving
- Cosmos World Foundation Model Platform for Physical AI
- Leveraging Edge Intelligence and LLMs to Advance 6G-Enabled Internet of Automated Defense Vehicles
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- Navigation World Models
- Revolutionizing QoE-Driven Network Management with Digital Agents in 6G
- Understanding World or Predicting Future? A Comprehensive Survey of World Models
- How Far is Video Generation from World Model: A Physical Law Perspective
- WorldSimBench: Towards Video Generation Models as World Simulators
- WALL-E: World Alignment by Rule Learning Improves World Model-based LLM Agents
- User-centric Immersive Communications in 6G: A Data-oriented Framework via Digital Twin
- Generative AI based Secure Wireless Sensing for ISAC Networks
- The Llama 3 Herd of Models
- Model Predictive Path Integral Control for Agile Unmanned Aerial Vehicles
- EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
- Toward Enhanced Reinforcement Learning-Based Resource Management via Digital Twin: Opportunities, Applications, and Challenges
- Evaluating the World Model Implicit in a Generative Model
- Generative AI for Deep Reinforcement Learning: Framework, Analysis, and Use Cases
- Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllability
- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond
- Artificial General Intelligence (AGI)-Native Wireless Systems: A Journey Beyond 6G
- WorldGPT: Empowering LLM as Multimodal World Model
- Digital Twin Assisted Intelligent Network Management for Vehicular Applications
- World Models for Autonomous Driving: An Initial Survey
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Learning and Leveraging World Models in Visual Representation Learning
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- SwiftCache: Model-Based Learning for Dynamic Content Caching in CDNs
- Generative AI for Secure Physical Layer Communications: A Survey
- Trustworthy Distributed AI Systems: Robustness, Privacy, and Governance
- Digital Twin-Based Network Management for Better QoE in Multicast Short Video Streaming
- A Survey of Recent Advances in Optimization Methods for Wireless Communications
- Small Dataset, Big Gains: Enhancing Reinforcement Learning by Offline Pre-Training with Model Based Augmentation
- World Models via Policy-Guided Trajectory Diffusion
- TWIST: Teacher-Student World Model Distillation for Efficient Sim-to-Real Transfer
- CQM: Curriculum Reinforcement Learning with a Quantized World Model
- MemGPT: Towards LLMs as Operating Systems
- GAIA-1: A Generative World Model for Autonomous Driving
- SafeDreamer: Safe Reinforcement Learning with World Models
- Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions
- Reasoning with Language Model is Planning with World Model
- PaLM 2 Technical Report
- Generative Pre-trained Transformer: A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions
- GPT (Generative Pre-Trained Transformer)— A Comprehensive Review on Enabling Technologies, Potential Applications, Emerging Challenges, and Future Directions
- Vision-Language Models for Vision Tasks: A Survey
- Vision-Language Models for Vision Tasks: A Survey
- GPT-4 Technical Report
- Transformer-based World Models Are Happy With 100k Interactions
- TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction
- LLaMA: Open and Efficient Foundation Language Models
- Joint Spectrum and Power Allocation for V2X Communications with Imperfect CSI
- Diffusion Models: A Comprehensive Survey of Methods and Applications
- Joint Optimization of Resource Allocation, Phase Shift and UAV Trajectory for Energy-Efficient RIS-Assisted UAV-Enabled MEC Systems
- A Generalist Agent
- DreamingV2: Reinforcement Learning with Discrete World Models without Reconstruction
- TransDreamer: Reinforcement Learning with Transformer World Models
- Path Planning for the Dynamic UAV-Aided Wireless Systems using Monte Carlo Tree Search
- DreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical Representations
- Generative Adversarial Networks
- World-GAN: a Generative Model for Minecraft Worlds
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Transformers in Vision: A Survey
- A Survey on Vision Transformer
- Mastering Atari with Discrete World Models
- Dreaming: Model-based Reinforcement Learning by Latent Imagination without Reconstruction
- Deep Reinforcement Learning Based Mode Selection and Resource Allocation for Cellular V2X Communications
- A Survey of Predictive Maintenance: Systems, Purposes and Approaches
- Dream to Control: Learning Behaviors by Latent Imagination
- Mastering Atari, Go, chess and shogi by planning with a learned model
- Asynchronous Methods for Model-Based Reinforcement Learning
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence with Edge Computing
- Edge Intelligence: Paving the Last Mile of Artificial Intelligence With Edge Computing
- Wireless Networks Design in the Era of Deep Learning: Model-Based,\n AI-Based, or Both?
- Wireless Networks Design in the Era of Deep Learning: Model-Based, AI-Based, or Both?
- Learning Latent Dynamics for Planning from Pixels
- Recurrent World Models Facilitate Policy Evolution
- Deep Reinforcement Learning based Resource Allocation for V2V Communications
- Deep Reinforcement Learning Based Resource Allocation for V2V Communications
- DeepMind Control Suite
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- Neural Discrete Representation Learning
- Generative Adversarial Networks: An Overview
- Attention Is All You Need
- Tutorial on Variational Autoencoders
- OpenAI Gym
Cited by
Related