Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
2025/07/01 by Haoxiang Luo, Luo, Haoxiang, Yinqiu Liu +17 · 9 citations
Computer Science · Engineering · #Big Data and Digital Economy #Multimodal Machine Learning Applications #Ferroelectric and Negative Capacitance Devices
paper · pdf · doi:10.48550/arxiv.2507.00672
Abstract
Edge computing enables real-time data processing closer to its source, thus improving the latency and performance of edge-enabled AI applications. However, traditional AI models often fall short when dealing with complex, dynamic tasks that require advanced reasoning and multimodal data processing. This survey explores the integration of multi-LLMs (Large Language Models) to address this in edge computing, where multiple specialized LLMs collaborate to enhance task performance and adaptability in resource-constrained environments. We review the transition from conventional edge AI models to single LLM deployment and, ultimately, to multi-LLM systems. The survey discusses enabling technologies such as dynamic orchestration, resource scheduling, and cross-domain knowledge transfer that are key for multi-LLM implementation. A central focus is on trusted multi-LLM systems, ensuring robust decision-making in environments where reliability and privacy are crucial. We also present multimodal multi-LLM architectures, where multiple LLMs specialize in handling different data modalities, such as text, images, and audio, by integrating their outputs for comprehensive analysis. Finally, we highlight future directions, including improving resource efficiency, trustworthy governance multi-LLM systems, while addressing privacy, trust, and robustness concerns. This survey provides a valuable reference for researchers and practitioners aiming to leverage multi-LLM systems in edge computing applications.
Citations
- Secure Multi-LLM Agentic AI and Agentification for Edge General Intelligence by Zero-Trust: A Survey
- Toward Edge General Intelligence with Agentic AI and Agentification: Concepts, Technologies, and Future Directions
- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges
- Mobility-Aware Asynchronous Federated Learning with Dynamic Sparsification
- Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
- TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems
- World Models for Cognitive Agents: Transforming Edge Intelligence in Future Networks
- Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception
- Chain-of-Thought for Large Language Model-empowered Wireless Communications
- Large Language Model-enhanced Reinforcement Learning for Low-Altitude Economy Networking
- Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control
- Temporal Spectrum Cartography in Low-Altitude Economy Networks: A Generative AI Framework with Multi-Agent Learning
- Memory-Efficient Split Federated Learning for LLM Fine-Tuning on Heterogeneous Mobile Devices
- Movable Antenna Enhanced Federated Fine-Tuning of Large Language Models via Hybrid Client Selection Optimization
- Model-Distributed Inference for Large Language Models at the Edge
- A Weighted Byzantine Fault Tolerance Consensus Driven Trusted Multiple Large Language Models Network
- A Trustworthy Multi-LLM Network: Challenges,Solutions, and A Use Case
- Edge Large AI Models: Collaborative Deployment and IoT Applications
- Covert Prompt Transmission for Secure Large Language Model Services
- Toward Realization of Low-Altitude Economy Networks: Core Architecture, Integrated Technologies, and Future Directions
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- Secure Physical Layer Communications for Low-Altitude Economy Networking: A Survey
- Large Language Models integration in Smart Grids
- EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices
- A Case Study of Scalable Content Annotation Using Multi-LLM Consensus and Human Review
- Improving the End-to-End Efficiency of Offline Inference for Multi-LLM Applications Based on Sampling and Simulation
- GameChat: Multi-LLM Dialogue for Safe, Agile, and Socially Optimal Multi-Agent Navigation in Constrained Environments
- Privacy-Enhancing Paradigms within Federated Multi-Agent Systems
- Wireless Hallucination in Generative AI-enabled Communications: Concepts, Issues, and Solutions
- Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
- LLM-Fusion: A Novel Multimodal Fusion Model for Accelerated Material Discovery
- FSPO: Few-Shot Preference Optimization of Synthetic Preference Data in LLMs Elicits Effective Personalization to Real Users
- Generative AI-enabled Wireless Communications for Robust Low-Altitude Economy Networking
- Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
- Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems
- A Tutorial on LLM Reasoning: Relevant Methods behind ChatGPT o1
- Cost-Saving LLM Cascades with Early Abstention
- When One LLM Drools, Multi-LLM Collaboration Rules
- On the Reasoning Capacity of AI Models and How to Quantify It
- Split Fine-Tuning for Large Language Models in Wireless Networks
- Multi-Agent Collaboration Mechanisms: A Survey of LLMs
- Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions
- Embodied AI-Enhanced Vehicular Networks: An Integrated Large Language Models and Reinforcement Learning Method
- Distributed Mixture-of-Agents for Edge Inference with Large Language Models
- Multi-LLM Text Summarization
- DRDST: Low-latency DAG Consensus through Robust Dynamic Sharding and Tree-broadcasting for IoV
- Multimodal Alignment and Fusion: A Survey
- TEESlice: Protecting Sensitive Neural Network Models in Trusted Execution Environments When Attackers have Pre-Trained Models
- Toward Democratized Generative AI in Next-Generation Mobile Edge Networks
- CE-CoLLM: Efficient and Adaptive Large Language Models Through Cloud-Edge Collaboration
- A Scalable Communication Protocol for Networks of Large Language Models
- TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
- A Review on Edge Large Language Models: Design, Execution, and Applications
- Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
- A Multi-LLM Debiasing Framework
- What is the Role of Small Models in the LLM Era: A Survey
- From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks
- GenAI-powered Multi-Agent Paradigm for Smart Urban Mobility: Opportunities and Challenges for Integrating Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) with Intelligent Transportation Systems
- Fine-Tuning and Deploying Large Language Models Over Edges: Issues and Approaches
- Convergence of Symbiotic Communications and Blockchain for Sustainable and Trustworthy 6G Wireless Networks
- Large Language Model (LLM)-enabled Graphs in Dynamic Networking
- Mobile Edge Intelligence for Large Language Models: A Contemporary Survey
- Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models
- Agreement-Based Cascading for Efficient Inference
- SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
- Dynamic Scheduling for Vehicle-to-Vehicle Communications Enhanced Federated Learning
- FedBiOT: LLM Local Fine-tuning in Federated Learning without Full Model
- EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
- Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
- Enhancing Supermarket Robot Interaction: A Multi-Level LLM Conversational Interface for Handling Diverse Customer Intents
- Leveraging Foundation Models for Multi-modal Federated Learning with Incomplete Modality
- UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
- Cost-Effective Online Multi-LLM Selection with Versatile Reward Models
- Optimizing Generative AI Networking: A Dual Perspective with Multi-Agent Systems and Mixture of Experts
- Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities
- Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities
- Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
- INMS: Memory Sharing for Large Language Model based Agents
- ProSecutor: Protecting Mobile AIGC Services on Two-Layer Blockchain via Reputation and Contract Theoretic Approaches
- Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
- Blockchain for Energy Market: A Comprehensive Survey
- On Protecting the Data Privacy of Large Language Models (LLMs): A Survey
- Levels of AI Agents: from Rules to Large Language Models
- Learning to Decode Collaboratively with Multiple Language Models
- Generative AI for Secure Physical Layer Communications: A Survey
- Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration
- Knowledge Fusion of Large Language Models
- SocraSynth: Multi-LLM Reasoning with Conditional Statistics
- An Edge-Cloud Collaboration Framework for Generative AI Service Provision with Synergetic Big Cloud Model and Small Edge Models
- Generative AI-driven Semantic Communication Networks: Architecture, Technologies and Applications
- Symbiotic Blockchain Consensus: Cognitive Backscatter Communications-enabled Wireless Blockchain Consensus
- Splitwise: Efficient generative LLM inference using phase splitting
- Large Language Models for Networking: Applications, Enabling Techniques, and Challenges
- ChatGPT and Other Large Language Models for Cybersecurity of Smart Grid Applications
- Split-and-Denoise: Protect large language model inference with local differential privacy
- Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities
- LLMCad: Fast and Scalable On-device Large Language Model Inference
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- In-context Autoencoder for Context Compression in a Large Language Model
- A Survey on Evaluation of Large Language Models
- A Survey on Evaluation of Large Language Models
- A Survey of Multimodal Information Fusion for Smart Healthcare: Mapping the Journey from Data to Wisdom
- Split Learning in 6G Edge Networks
- Domain Specialization as the Key to Make Large Language Models Disruptive: A Comprehensive Survey
- ESIA: An Efficient and Stable Identity Authentication for Internet of Vehicles
- LLM-Pruner: On the Structural Pruning of Large Language Models
- Towards Building the Federated GPT: Federated Instruction Tuning
- ESCM: An Efficient and Secure Communication Mechanism for UAV Networks
- Generative AI-enabled Vehicular Networks: Fundamentals, Framework, and Case Study
- Data-centric Artificial Intelligence: A Survey
- The pipeline for the continuous development of artificial intelligence models—Current state of research and practice
- Distill-VQ: Learning Retrieval Oriented Vector Quantization By Distilling Knowledge from Dense Embeddings
- EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- TinyBERT: Distilling BERT for Natural Language Understanding
- Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training
- Confidential Prompting: Privacy-preserving LLM Inference on Cloud
Cited by
Related