Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
2025/06/12 by Zhiyang Xu, Jiuhai Chen, Xu, Zhiyang +23 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2506.10395
openalex publication_date 2025/06/12 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Recent advances in large language models (LLMs) have enabled multimodal foundation models to tackle both image understanding and generation within a unified framework. Despite these gains, unified models often underperform compared to specialized models in either task. A key challenge in developing unified models lies in the inherent differences between the visual features needed for image understanding versus generation, as well as the distinct training processes required for each modality. In this work, we introduce Pisces, an auto-regressive multimodal foundation model that addresses this challenge through a novel decoupled visual encoding architecture and tailored training techniques optimized for multimodal generation. Combined with meticulous data curation, pretraining, and finetuning, Pisces achieves competitive performance in both image understanding and image generation. We evaluate Pisces on over 20 public benchmarks for image understanding, where it demonstrates strong performance across a wide range of tasks. Additionally, on GenEval, a widely adopted benchmark for image generation, Pisces exhibits robust generative capabilities. Our extensive analysis reveals the synergistic relationship between image understanding and generation, and the benefits of using separate visual encoders, advancing the field of unified multimodal models.
Citations
- Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
- LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- Multimodal Instruction Tuning with Conditional Mixture of LoRA
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning
- MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- Generative Multimodal Models are In-Context Learners
- VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation
- CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- Improved Baselines with Visual Instruction Tuning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- PixArt-α: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
- DreamLLM: Synergistic Multimodal Comprehension and Creation
- NExT-GPT: Any-to-Any Multimodal LLM
- Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization
- Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- MMBench: Is Your Multi-modal Model an All-around Player?
- Emu: Generative Pretraining in Multimodality
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration
- LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning
- Generating Images with Multimodal Language Models
- Any-to-Any Generation via Composable Diffusion
- Evaluating Object Hallucination in Large Vision-Language Models
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models
- Visual Instruction Tuning
- Sigmoid Loss for Language Image Pre-Training
- EVA-CLIP: Improved Training Techniques for CLIP at Scale
- LLaMA: Open and Efficient Foundation Language Models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction Tuning
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
- Flamingo: a Visual Language Model for Few-Shot Learning
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
- CM3: A Causal Masked Multimodal Model of the Internet
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Masked Autoencoders Are Scalable Vision Learners
- InfographicVQA
- Learning Transferable Visual Models From Natural Language Supervision
- Taming Transformers for High-Resolution Image Synthesis
- DocVQA: A Dataset for VQA on Document Images
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Towards VQA Models That Can Read
- VizWiz Grand Challenge: Answering Visual Questions from Blind People
- Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
- A Diagram Is Worth A Dozen Images
- Microsoft COCO: Common Objects in Context
- Retrieval-Augmented Multimodal Language Modeling
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Cited by
Related