Vision as Unified Multimodal Generation
2026/07/07 by Xiaoyang Han, Jianhua Li, Kewang Deng +14 · 1 voice
Computer Science · #cs.CV
paper · pdf
Abstract
We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task-specific architectures. Under this formulation, SenseNova-Vision uses natural-language instructions and optional visual prompts to specify tasks, target regions or views, and decoding conventions, and generates responses as text for symbolic outputs, images for dense spatial predictions, or mixed text-and-image outputs for compositional tasks. To support large-scale training, we convert diverse computer vision annotations into instruction-response examples compatible with these generation spaces, resulting in the SenseNova-Vision Corpus, a computer-vision instruction-response corpus spanning text, image, and mixed targets. Starting from an off-the-shelf pretrained unified multimodal model, SenseNova-Vision is trained primarily on this corpus, with auxiliary multimodal data used as a capability-preserving mixture, and requires no task-specific prediction heads or architectural modifications. The resulting model covers a broad range of vision tasks, including detection, OCR, keypoint estimation, segmentation, depth estimation, surface normal prediction, point maps, and camera pose estimation, while supporting language-defined variants that combine category, color, region, and other visual cues. Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry. These results suggest unified multimodal generation as a scalable route for integrating computer vision capabilities into general-purpose foundation models. The model and corpus are publicly available.
Citations
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- Image Generators are Generalist Vision Learners
- Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
- Masked Depth Modeling for Spatial Perception
- nuScenes Revisited: Progress and Challenges in Autonomous Driving
- Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
- Qwen3-VL Technical Report
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- ConsistCompose: Unified Multimodal Layout Control for Image Composition
- Depth Anything 3: Recovering the Visual Space from Any Views
- Visual Bridge: Universal Visual Perception Representations Generating
- Detect Anything via Next Point Prediction
- MapAnything: Universal Feed-Forward Metric 3D Reconstruction
- From Editor to Dense Geometry Estimator
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- X-SAM: From Segment Anything to Any Segmentation
- Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
- MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details
- Emerging Properties in Unified Multimodal Pretraining
- VGGT: Visual Geometry Grounded Transformer
- DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- MegaSynth: Scaling Up 3D Scene Reconstruction with Synthesized Data
- ChatRex: Taming Multimodal LLM for Joint Perception and Understanding
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent
- OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
- Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
- Text4Seg: Reimagining Image Segmentation as Text Generation
- Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
- EAFormer: Scene Text Segmentation with Edge-Aware Transformers
- Depth Anything V2
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
- Chameleon: Mixed-Modal Early-Fusion Foundation Models
- COCONut: Modernizing COCO Segmentation
- PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model
- SceneScript: Reconstructing Scenes With An Autoregressive Structured Language Model
- Rethinking Inductive Biases for Surface Normal Estimation
- RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D Videos
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
- LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model
- DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision
- APTv2: Benchmarking Animal Pose Estimation and Tracking with a Large-scale Dataset and Beyond
- InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
- DUSt3R: Geometric 3D Vision Made Easy
- Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
- PixelLM: Pixel Reasoning with Large Multimodal Model
- HumanRef: Single Image to 3D Human Generation via Reference-Guided Diffusion
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks
- Language-guided Robot Grasping: CLIP-based Referring Grasp Synthesis in Clutter
- GLaMM: Pixel Grounding Large Multimodal Model
- GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding
- ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes
- LISA: Reasoning Segmentation via Large Language Model
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- GRES: Generalized Referring Expression Segmentation
- ICDAR 2023 Competition on Hierarchical Text Detection and Recognition
- Visual Instruction Tuning
- DINOv2: Learning Robust Visual Features without Supervision
- V3Det: Vast Vocabulary Visual Detection Dataset
- Micrograph segmentations for DDEVD
- Segment Anything
- Annotated Point Clouds and Images from a Spatial-Semantic Perception Pipeline: Robotized Deconstruction at the Reference Construction Site, Aachen
- Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial Scenes
- OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation
- PACO: Parts and Attributes of Common Objects
- Objaverse: A Universe of Annotated 3D Objects
- Images Speak in Images: A Generalist Painter for In-Context Visual Learning
- PIDray: A Large-scale X-ray Benchmark for Real-World Prohibited Item Detection
- PIDray: A Large-Scale X-ray Benchmark for Real-World Prohibited Item Detection
- Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
- High-Quality Entity Segmentation
- Learning-based Inverse Rendering of Complex Indoor Scenes with Differentiable Monte Carlo Raytracing
- VizWiz-FewShot: Locating Objects in Images Taken by People With Visual Impairments
- Panoptic Scene Graph Generation
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- A Unified Sequence Interface for Vision Tasks
- APT-36K: A Large-scale Benchmark for Animal Pose Estimation and Tracking
- Flamingo: a Visual Language Model for Few-Shot Learning
- Towards End-to-End Unified Scene Text Detection and Layout Analysis
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- Masked-attention Mask Transformer for Universal Image Segmentation
- PartImageNet: A Large, High-Quality Dataset of Parts
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling
- Masked Autoencoders Are Scalable Vision Learners
- Pix2seq: A Language Modeling Framework for Object Detection
- Panoptic nuScenes: A Large-Scale Benchmark for LiDAR Panoptic Segmentation and Tracking
- AP-10K: A Benchmark for Animal Pose Estimation in the Wild
- ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes
- TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
- Human-centric Relation Segmentation: Dataset and Solution
- RaidaR: A Rich Annotated Image Dataset of Rainy Street Scenes
- Spatial Dual-Modality Graph Reasoning for Key Information Extraction
- FAIR1M: A Benchmark Dataset for Fine-grained Object Recognition in High-Resolution Remote Sensing Imagery
- STEP: Segmenting and Tracking Every Pixel
- A Fine-Grained Dataset and its Efficient Semantic Segmentation for\n Unstructured Driving Scenarios
- Rethinking Text Segmentation: A Novel Dataset and A Text-Specific Refinement Approach
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding
- TTPLA: An Aerial-Image Dataset for Detection and Segmentation of Transmission Towers and Power Lines
- TrashCan: A Semantically-Segmented Dataset towards Visual Detection of Marine Debris
- Denoising Diffusion Probabilistic Models
- Language Models are Few-Shot Learners
- End-to-End Object Detection with Transformers
- SCRDet++: Detecting Small, Cluttered and Rotated Objects via Instance-Level Feature Denoising and Rotation Loss Smoothing
- Fashionpedia: Ontology, Segmentation, and an Attribute Localization Dataset
- Semantic Segmentation of Underwater Imagery: Dataset and Benchmark
- Segmenting Transparent Objects in the Wild
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- RAFT: Recurrent All-Pairs Field Transforms for Optical Flow
- Rethinking Object Detection in Retail Stores
- Detection and Tracking Meet Drones Challenge
- Scale Match for Tiny Person Detection
- BlendedMVS: A Large-scale Dataset for Generalized Multi-view Stereo Networks
- WiderPerson: A Diverse Dataset for Dense Pedestrian Detection in the Wild
- PST900: RGB-Thermal Calibration, Dataset and Segmentation Network
- A Real-Time Cross-modality Correlation Filtering Method for Referring Expression Comprehension
- ICDAR 2019 Robust Reading Challenge on Reading Chinese Text on Signboard
- PubLayNet: largest dataset ever for document layout analysis
- DIODE: A Dense Indoor and Outdoor DEpth Dataset
- ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition -- RRC-MLT-2019
- LVIS: A Dataset for Large Vocabulary Instance Segmentation
- The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation
- DeepFashion2: A Versatile Benchmark for Detection, Pose Estimation, Segmentation and Re-Identification of Clothing Images
- A Hierarchical Grocery Store Image Dataset with Visual and Semantic Labels
- CrowdPose: Efficient Crowded Scenes Pose Estimation and A New Benchmark
- Detector-in-Detector: Multi-Level Analysis for Human-Parts
- IDD: A Dataset for Exploring Problems of Autonomous Navigation in Unconstrained Environments
- YouTube-VOS: A Large-Scale Video Object Segmentation Benchmark
- Instance-level Human Parsing via Part Grouping Network
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning
- Evaluation of CNN-based Single-Image Depth Estimation Methods
- CrowdHuman: A Benchmark for Detecting Human in a Crowd
- MVTec D2S: Densely Segmented Supermarket Dataset
- Taskonomy: Disentangling Task Transfer Learning
- Falling Things: A Synthetic Dataset for 3D Object Detection and Pose Estimation
- DeepMVS: Learning Multi-view Stereopsis
- Analysis of Hand Segmentation in the Wild
- ICDAR2017 Competition on Reading Chinese Text in the Wild (RCTW-17)
- Drone-based Object Counting by Spatially Regularized Regional Proposal Network
- Mask R-CNN
- Look into Person: Self-supervised Structure-sensitive Learning and A New Benchmark for Human Parsing
- ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- Modeling Context in Referring Expressions
- Virtual Worlds as Proxy for Multi-Object Tracking Analysis
- Synthetic Data for Text Localisation in Natural Images
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
- Microsoft COCO: Common Objects in Context
- Vision meets robotics: The KITTI dataset
- Open-world Text-specified Object Counting
- Scaling Out-of-Distribution Detection for Real-World Settings
Discussions
Related