FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
2025/12/14 by Yue Jiang, Jiang, Yue, Dingkang Yang +15 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Speech and dialogue systems #Topic Modeling
paper · pdf · doi:10.48550/arxiv.2512.12756
openalex publication_date 2025/12/14 · openalex created_date 2025/12/17 · openalex updated_date 2026/07/28
Abstract
Despite rapid progress in multimodal large language models (MLLMs) and emerging omni-modal architectures, current benchmarks remain limited in scope and integration, suffering from incomplete modality coverage, restricted interaction to text-centric outputs, and weak interdependence and complementarity among modalities. To bridge these gaps, we introduce FysicsWorld, the first unified full-modality benchmark that supports bidirectional input-output across image, video, audio, and text, enabling comprehensive any-to-any evaluation across understanding, generation, and reasoning. FysicsWorld encompasses 16 primary tasks and 3,268 curated samples, aggregated from over 40 high-quality sources and covering a rich set of open-domain categories with diverse question types. We also propose the Cross-Modal Complementarity Screening (CMCS) strategy integrated in a systematic data construction framework that produces omni-modal data for spoken interaction and fusion-dependent cross-modal reasoning. Through a comprehensive evaluation of over 30 state-of-the-art baselines, spanning MLLMs, modality-specific models, unified understanding-generation models, and omni-modal language models, FysicsWorld exposes the performance disparities and limitations across models in understanding, generation, and reasoning. Our benchmark establishes a unified foundation and strong baselines for evaluating and advancing next-generation full-modality architectures.
Citations
- SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension
- Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection
- Emu3.5: Native Multimodal Models are World Learners
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- HunyuanImage 3.0 Technical Report
- Seedream 4.0: Toward Next-generation Multimodal Image Generation
- Qwen3-Omni Technical Report
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- Qwen-Image Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Ovis-U1 Technical Report
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Show-o2: Improved Native Unified Multimodal Models
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- Seedance 1.0: Exploring the Boundaries of Video Generation Models
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
- SeedEdit 3.0: Fast and High-Quality Generative Image Editing
- DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
- Emerging Properties in Unified Multimodal Pretraining
- MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
- Step1X-Edit: A Practical Framework for General Image Editing
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Video-Bench: Human-Aligned Video Generation Benchmark
- OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
- Qwen2.5-Omni Technical Report
- Gemma 3 Technical Report
- WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
- Qwen2.5-VL Technical Report
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Baichuan-Omni-1.5 Technical Report
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
- AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
- LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
- MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
- AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
- MedAide: Information Fusion and Anatomy of Medical Intents via LLM-based Agent Collaboration
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- Qwen2-Audio Technical Report
- Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
- Towards Context-Aware Emotion Recognition Debiasing from a Causal Demystification Perspective via De-confounded Training
- Towards Context-Aware Emotion Recognition Debiasing From a Causal Demystification Perspective via De-Confounded Training
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications
- Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
- AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
- VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
- SALMONN: Towards Generic Hearing Abilities for Large Language Models
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
- MMBench: Is Your Multi-modal Model an All-around Player?
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Visual Instruction Tuning
- GPT-4 Technical Report
- BERTScore: Evaluating Text Generation with BERT
Cited by
Related