Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
2025/08/07 by Feng, Kunyu, Ma, Yue, Zhang, Xinhua +9 · 8 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2508.05580
Abstract
With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real-world data remains costly and time-consuming, hindering the development of downstream applications. While some works attempt to collect task-specific data via a rendering process, most approaches still rely on manual scene construction, limiting their scalability and accuracy. To address these challenges, we propose Follow-Your-Instruction, a Multimodal Large Language Model (MLLM)-driven framework for automatically synthesizing high-quality 2D, 3D, and 4D data. Our Follow-Your-Instruction first collects assets and their associated descriptions through multimodal inputs using the MLLM-Collector. Then it constructs 3D layouts, and leverages Vision-Language Models (VLMs) for semantic refinement through multi-view scenes with the MLLM-Generator and MLLM-Optimizer, respectively. Finally, it uses MLLM-Planner to generate temporally coherent future frames. We evaluate the quality of the generated data through comprehensive experiments on the 2D, 3D, and 4D generative tasks. The results show that our synthetic data significantly boosts the performance of existing baseline models, demonstrating Follow-Your-Instruction's potential as a scalable and effective data engine for generative intelligence.
Citations
- Controllable Video Generation: A Survey
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
- Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- Gemma 3 Technical Report
- Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- EEdit: Rethinking the Spatial and Temporal Redundancy for Efficient Image Editing
- TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
- Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate Gradients
- Qwen2.5-VL Technical Report
- Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
- SmartEraser: Remove Anything from Images using Masked-Region Guidance
- RORem: Training a Robust Object Remover with Human-in-the-Loop
- Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
- ReCap: Better Gaussian Relighting with Cross-Environment Captures
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
- MV-Adapter: Multi-view Consistent Image Generation Made Easy
- InstantSwap: Fast Customized Concept Swapping across Sharp Shape Differences
- Evaluating Text-to-Image Diffusion Models for Texturing Synthetic Data
- Taming Rectified Flow for Inversion and Editing
- DiT4Edit: Diffusion Transformer for Image Editing
- MLLM as Retriever: Interactively Learning Multimodal Retrieval for Embodied Agents
- Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
- VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
- RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
- COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video Editing
- Towards Multiple Character Image Animation Through Enhancing Implicit Decoupling
- Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation
- Ovis: Structural Embedding Alignment for Multimodal Large Language Model
- MultiBooth: Towards Generating All Your Concepts in an Image from Text
- Towards Realistic Scene Generation with LiDAR Diffusion Models
- SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
- MagicStick: Controllable Video Editing via Control Handle Transformations
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- Clarity ChatGPT: An Interactive and Adaptive Processing System for Image Restoration and Enhancement
- Synthetic Data Generation for Bridging Sim2Real Gap in a Production Environment
- Exploring Limits of Diffusion-Synthetic Training with Weakly Supervised Semantic Segmentation
- AgentBench: Evaluating LLMs as Agents
- Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
- Video-P2P: Video Editing with Cross-attention Control
- Scalable Diffusion Models with Transformers
- LAION-5B: An open large-scale dataset for training next generation image-text models
- High-Resolution Image Synthesis with Latent Diffusion Models
- High-Resolution Image Synthesis with Latent Diffusion Models
- LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
- Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval
- Learning Transferable Visual Models From Natural Language Supervision
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related