NExT-GPT: Any-to-Any Multimodal LLM
2023/09/11 by Shengqiong Wu, Wu, Shengqiong, Fei, Hao +6 · 104 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2309.05519
openalex publication_date 2023/09/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities. As we humans always perceive the world and communicate with people through various modalities, developing any-to-any MM-LLMs capable of accepting and delivering content in any modality becomes essential to human-level AI. To fill the gap, we present an end-to-end general-purpose any-to-any MM-LLM system, NExT-GPT. We connect an LLM with multimodal adaptors and different diffusion decoders, enabling NExT-GPT to perceive inputs and generate outputs in arbitrary combinations of text, images, videos, and audio. By leveraging the existing well-trained highly-performing encoders and decoders, NExT-GPT is tuned with only a small amount of parameter (1%) of certain projection layers, which not only benefits low-cost training and also facilitates convenient expansion to more potential modalities. Moreover, we introduce a modality-switching instruction tuning (MosIT) and manually curate a high-quality dataset for MosIT, based on which NExT-GPT is empowered with complex cross-modal semantic understanding and content generation. Overall, our research showcases the promising possibility of building an AI agent capable of modeling universal modalities, paving the way for more human-like AI research in the community. Project page: https://next-gpt.github.io/
Cited by
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- Video Understanding: From Geometry and Semantics to Unified Models
- Structured Prompting and LLM Ensembling for Multimodal Conversational Aspect-based Sentiment Analysis
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- Sheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site
- Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- FlowBind: Efficient Any-to-Any Generation with Bidirectional Flows
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- ChronusOmni: Improving Time Awareness of Omni Large Language Models
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
- PhotoFramer: Multi-modal Image Composition Instruction
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
- SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- QuAnTS: Question Answering on Time Series
- NAPS: Attention-Based Fusion of Heterogeneous Physiological Signals
- ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
- Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Large Emotional World Model
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Robust Distortion-Free Watermark for Autoregressive Audio Generation Models
- UniMedVL: Unifying Medical Multimodal Understanding And Generation Through Observation-Knowledge-Analysis
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- End-to-End Multi-Modal Diffusion Mamba
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- A Survey on Agentic Multimodal Large Language Models
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Towards Reliable LLM-based Robot Planning via Combined Uncertainty Estimation
- VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- MuSLR: Multimodal Symbolic Logical Reasoning
- MotionVerse: A Unified Multimodal Framework for Motion Comprehension, Generation and Editing
- AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
- Planning with Unified Multimodal Models
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- VIG-RL: Learning to Search and Insert for Verified Image Grounding
- COLT: Enhancing Video Large Language Models with Continual Tool Usage
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
- Revealing Multimodal Causality with Large Language Models
- UnifiedVisual: A Framework for Constructing Unified Vision-Language Datasets
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- Testing chatbots on the creation of encoders for audio conditioned image generation
- Interleaving Reasoning for Better Text-to-Image Generation
- Effectively obtaining acoustic, visual and textual data from videos
- AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
- Audio-Guided Visual Editing with Complex Multi-Modal Prompts
- AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation
- MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
- Towards Scalable and Interpretable Mobile App Risk Analysis via Large Language Models
- Cryfish: On deep audio analysis with Large Language Models
- E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
- Agentic Design Review System
- A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- A Comprehensive Review of Datasets for Clinical Mental Health AI Systems
- Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education
- Never compromise with vulnerabilities: a comprehensive survey on AI governance
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
- VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
- Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- Advancing Compositional LLM Reasoning with Structured Task Relations in Interactive Multimodal Communications
- MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
- Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
- DatasetAgent: A Novel Multi-Agent System for Auto-Constructing Datasets from Real-World Images
- Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs
- NeoBabel: A Multilingual Open Tower for Visual Generation
- UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification
- MotionGPT3: Human Motion as a Second Modality
- Where, What, Why: Towards Explainable Driver Attention Prediction
- Are Large Language Models Capable of Deep Relational Reasoning? Insights from DeepSeek-R1 and Benchmark Comparisons
- XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge
- ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- ReactEMG: Stable, Low-Latency Intent Detection from sEMG via Masked Modeling
- MATE: LLM-Powered Multi-Agent Translation Environment for Accessibility Applications
- Let Your Video Listen to Your Music!
- Federated Learning from Molecules to Processes: A Perspective
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Related