WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
2025/03/10 by Niu, Yuwei, Ning, Munan, Zheng, Mengren +8 · 95 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #I.2.10 #I.2.7 #I.4.9
paper · doi:10.48550/arxiv.2503.07265
Abstract
Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evaluation standards predominantly focus on image realism and shallow text-image alignment, lacking a comprehensive assessment of complex semantic understanding and world knowledge integration in text-to-image generation. To address this challenge, we propose WISE, the first benchmark specifically designed for World Knowledge-Informed Semantic Evaluation. WISE moves beyond simple word-pixel mapping by challenging models with 1000 meticulously crafted prompts across 25 subdomains in cultural common sense, spatio-temporal reasoning, and natural science. To overcome the limitations of traditional CLIP metric, we introduce WiScore, a novel quantitative metric for assessing knowledge-image alignment. Through comprehensive testing of 20 models (10 dedicated T2I models and 10 unified multimodal models) using 1,000 structured prompts spanning 25 subdomains, our findings reveal significant limitations in their ability to effectively integrate and apply world knowledge during image generation, highlighting critical pathways for enhancing knowledge incorporation and application in next-generation T2I models. Code and data are available at \hrefhttps://github.com/PKU-YuanGroup/WISEPKU-YuanGroup/WISE.
Cited by
- LongCat-Image Technical Report
- ThinkGen: Generalized Thinking for Visual Generation
- UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning
- EditThinker: Unlocking Iterative Reasoning for Any Image Editor
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- Understanding and Harnessing Sparsity in Unified Multimodal Models
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
- Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
- ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- UniREditBench: A Unified Reasoning-based Image Editing Benchmark
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- PairUni: Pairwise Training for Unified Multimodal Language Models
- Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- Open Multimodal Retrieval-Augmented Factual Image Generation
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- PICABench: How Far Are We from Physically Realistic Image Editing?
- Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- World-To-Image: Grounding Text-to-Image Generation with Agent-Driven World Knowledge
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
- IRIS: Intrinsic Reward Image Synthesis
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- UniECG: Understanding and Generating ECG in One Unified Model
- MEF: A Systematic Evaluation Framework for Text-to-Image Models
- Remote Sensing-Oriented World Model
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- GenExam: A Multidisciplinary Text-to-Image Exam
- Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration
- Unified Multimodal Model as Auto-Encoder
- Reconstruction Alignment Improves Unified Multimodal Models
- Interleaving Reasoning for Better Text-to-Image Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- A Survey of Multimodal Hallucination Evaluation and Detection
- T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
- Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- Show-o2: Improved Native Unified Multimodal Models
- MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
- UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
- TIIF-Bench: How Does Your T2I Model Follow Your Instructions?
- GenSpace: Benchmarking Spatially-Aware Image Generation
- R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks
- ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
- MMaDA: Multimodal Large Diffusion Language Models
- Emerging Properties in Unified Multimodal Pretraining
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- InterleaveThinker: Reinforcing Agentic Interleaved Generation
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
- TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training
- LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
- ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability
Related