OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
2024/11/11 by Cong Wei, Wei, Cong, Zhiqiang Xiong +9 · 47 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #AI in cancer detection #Artificial Intelligence (cs.AI) #Biomedical Text Mining and Ontologies #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques
paper · pdf · doi:10.48550/arxiv.2411.07199
openalex publication_date 2024/11/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Instruction-guided image editing methods have demonstrated significant potential by training diffusion models on automatically synthesized or manually annotated image editing pairs. However, these methods remain far from practical, real-life applications. We identify three primary challenges contributing to this gap. Firstly, existing models have limited editing skills due to the biased synthesis process. Secondly, these methods are trained with datasets with a high volume of noise and artifacts. This is due to the application of simple filtering methods like CLIP-score. Thirdly, all these datasets are restricted to a single low resolution and fixed aspect ratio, limiting the versatility to handle real-world use cases. In this paper, we present \omniedit, which is an omnipotent editor to handle seven different image editing tasks with any aspect ratio seamlessly. Our contribution is in four folds: (1) \omniedit is trained by utilizing the supervision from seven different specialist models to ensure task coverage. (2) we utilize importance sampling based on the scores provided by large multimodal models (like GPT-4o) instead of CLIP-score to improve the data quality. (3) we propose a new editing architecture called EditNet to greatly boost the editing success rate, (4) we provide images with different aspect ratios to ensure that our model can handle any image in the wild. We have curated a test set containing images of different aspect ratios, accompanied by diverse instructions to cover different tasks. Both automatic evaluation and human evaluations demonstrate that \omniedit can significantly outperform all the existing models. Our code, dataset and model will be available at https://tiger-ai-lab.github.io/OmniEdit/
Cited by
- LongCat-Image Technical Report
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Region-Constraint In-Context Generation for Instructional Video Editing
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- EasyV2V: A High-quality Instruction-based Video Editing Framework
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- Refaçade: Editing Object with Given Reference Texture
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
- Instruction-based image editing: a survey on data, models, evaluation, and applications
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
- OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- ReMix: Towards a Unified View of Consistent Character Generation and Editing
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- DreamOmni2: Multimodal Instruction-based Editing and Generation
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- TBStar-Edit: From Image Editing Pattern Shifting to Consistency Enhancement
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- SpotEdit: Evaluating Visually-Guided Image Editing Methods
- Visual Autoregressive Modeling for Instruction-Guided Image Editing
- X2Edit: Revisiting Arbitrary-Instruction Image Editing through Self-Constructed Data and Task-Aware Representation Learning
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- DreamVE: Unified Instruction-based Image and Video Editing
- The Promise of RL for Autoregressive Image Editing
- Trade-offs in Image Generation: How Do Different Dimensions Interact?
- GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
Related