MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
2026/07/28 by Mingqiao Ye, Zhaochong An, Zhitong Gao +11
#cs.CV #cs.AI #cs.LG
paper · pdf
Abstract
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly in scientific domains such as ecology and astronomy. Existing any-to-any models are typically trained from scratch using encoder-decoder or diffusion architectures, impacting their performance and preventing them from using strong pre-trained decoder-only models as a prior. In this work, we investigate decoder-only any-to-any multimodal modeling, which treats all modalities symmetrically and supports arbitrary modalities as inputs and outputs without modality-specific heads, losses, or task pipelines. Because every modality is both an input and an output of the same model, the resulting model, named Modus, can support a range of applications, such as chained generation through intermediate modalities or cross-modal self-verification by scoring the model's own outputs with another generated modality. Modus demonstrates strong out-of-the-box performance and is competitive with specialist and multitask baselines using a single model across various benchmarks. All materials are open-sourced at https://modus-multimodal.epfl.ch/.
Citations
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- HunyuanImage 3.0 Technical Report
- InfGen: A Resolution-Agnostic Paradigm for Scalable Image Synthesis
- TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Inference-time Scaling of Diffusion Models through Classical Search
- Emerging Properties in Unified Multimodal Pretraining
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
- BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
- Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
- A General Framework for Inference-time Scaling and Steering of Diffusion Models
- DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
- Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
- One Diffusion to Generate Them All
- Human Motion Instruction Tuning
- JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
- GPT-4o System Card
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Emu3: Next-Token Prediction is All You Need
- Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- LLaVA-OneVision: Easy Visual Task Transfer
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Depth Anything V2
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image
- Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
- AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
- Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- 4M: Massively Multimodal Masked Modeling
- Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
- GLaMM: Pixel Grounding Large Multimodal Model
- Qwen Technical Report
- NExT-GPT: Any-to-Any Multimodal LLM
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- ImageBind: One Embedding Space To Bind Them All
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- Micrograph segmentations for DDEVD
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Flow Matching for Generative Modeling
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- Exploring Plain Vision Transformer Backbones for Object Detection
- 3D Common Corruptions and Data Augmentation
- Learning Transferable Visual Models From Natural Language Supervision
Related