GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
2025/07/01 by GLM-V Team, V Team, : +179 · 1 voice · 186 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CV #cs.LG
paper · pdf · doi:10.48550/arxiv.2507.01006
arxiv published 2025/07/01 · arxiv updated 2026/01/01
Abstract
We present GLM-4.1V-Thinking, GLM-4.5V, and GLM-4.6V, a family of vision-language models (VLMs) designed to advance general-purpose multimodal understanding and reasoning. In this report, we share our key findings in the development of the reasoning-centric training framework. We first develop a capable vision foundation model with significant potential through large-scale pre-training, which arguably sets the upper bound for the final performance. We then propose Reinforcement Learning with Curriculum Sampling (RLCS) to unlock the full potential of the model, leading to comprehensive capability enhancement across a diverse range of tasks, including STEM problem solving, video understanding, content recognition, coding, grounding, GUI-based agents, and long document interpretation. In a comprehensive evaluation across 42 public benchmarks, GLM-4.5V achieves state-of-the-art performance on nearly all tasks among open-source models of similar size, and demonstrates competitive or even superior results compared to closed-source models such as Gemini-2.5-Flash on challenging tasks including Coding and GUI Agents. Meanwhile, the smaller GLM-4.1V-9B-Thinking remains highly competitive-achieving superior results to the much larger Qwen2.5-VL-72B on 29 benchmarks. We open-source both GLM-4.1V-9B-Thinking and GLM-4.5V. We further introduce the GLM-4.6V series, open-source multimodal models with native tool use and a 128K context window. A brief overview is available at https://z.ai/blog/glm-4.6v. Code, models and more information are released at https://github.com/zai-org/GLM-V.
Citations
Cited by
- From Generated Human Videos to Physically Plausible Robot Trajectories
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
- DECEPTICON: How Dark Patterns Manipulate Web Agents
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
- HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research
- An LMM for Precisely Grounding Elements in Documents
- VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
- Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- Streaming Video Instruction Tuning
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- MapTrace: Scalable Data Generation for Route Tracing on Maps
- M3-Verse: A "Spot the Difference" Challenge for Large Multimodal Models
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- TimeSeries2Report prompting enables adaptive large language model management of lithium-ion batteries
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
- Tracing the ongoing emergence of human-like reasoning in Large Language Models
- 6 Fingers, 1 Kidney: Natural Adversarial Medical Images Reveal Critical Weaknesses of Vision-Language Models
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- YingVideo-MV: Music-Driven Multi-Stage Video Generation
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models
- ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
- Med-CRAFT: Automated Construction of Interpretable and Multi-Hop Video Workloads via Knowledge Graph Traversal
- RealAppliance: Let High-fidelity Appliance Assets Controllable and Workable as Aligned Real Manuals
- Thinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn Interaction
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Geometrically-Constrained Agent for Spatial Reasoning
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- Understanding the Effects of Distractors on Reasoning Vision-Language Models
- CaptionQA: Is Your Caption as Useful as the Image Itself?
- GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- Fara-7B: An Efficient Agentic Model for Computer Use
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables
- P1: Mastering Physics Olympiads with Reinforcement Learning
- Video Finetuning Improves Reasoning Between Frames
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
- UI2CodeN: A Visual Language Model for Test-Time Scalable Interactive UI-to-Code Generation
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
- V-Thinker: Interactive Thinking with Images
- NVIDIA Nemotron Nano V2 VL
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- PreferThinker: Reasoning-based Personalized Image Preference Assessment
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
- CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
- Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
- SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
- DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
- LightAgent: Mobile Agentic Foundation Models
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
- CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
- See or Say Graphs: Agent-Driven Scalable Graph Structure Understanding with Vision-Language Models
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- The Art of Scaling Reinforcement Learning Compute for LLMs
- MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- SVAG-Bench: A Large-Scale Benchmark for Multi-Instance Spatio-temporal Video Action Grounding
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- A Survey on Agentic Multimodal Large Language Models
- Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
- GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- Self-signals Driven Multi-LLM Debate for Efficient and Accurate Reasoning
- EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- MMA-ASIA: A Multilingual and Multimodal Alignment Framework for Culturally-Grounded Evaluation
- Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- WebRenderBench: Enhancing Web Interface Generation through Layout-Style Consistency and Reinforcement Learning
- SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- Large Language Models Can Perform Automatic Modulation Classification via Discretized Self-supervised Candidate Retrieval
- TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs
- MR2-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- VideoScore2: Think before You Score in Generative Video Evaluation
- GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformation
- Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
- Score the Steps, Not Just the Goal: VLM-Based Subgoal Evaluation for Robotic Manipulation
- PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- OpenGVL -- Benchmarking Visual Temporal Progress for Data Curation
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- OpenLens AI: Fully Autonomous Research Agent for Health Infomatics
- GenExam: A Multidisciplinary Text-to-Image Exam
- SAIL-VL2 Technical Report
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Measuring Epistemic Humility in Multimodal Large Language Models
- Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
- RoboMatch: A Unified Mobile-Manipulation Teleoperation Platform with Auto-Matching Network Architecture for Long-Horizon Tasks
- VL Norm: Rethink Loss Aggregation in RLVR
- VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents
- Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
- PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- Baichuan-M2: Scaling Medical Capability with Large Verifier System
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- Intern-S1: A Scientific Multimodal Foundation Model
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
- ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- VASPilot: MCP-Facilitated Multi-Agent Intelligence for Autonomous VASP Simulations
- GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines
- IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?
- Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
- SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy
- Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
- Semantic Frame Interpolation
- TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
- SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- Towards Multimodal Graph Large Language Model
- Can Vision Language Models Infer Human Gaze Direction? A Controlled Study
- LUT: Latent Utility Training for Visual Reasoning
- SPR-128K: A New Benchmark for Spatial Plausibility Reasoning with Multimodal Large Language Models
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Prior Reinforce: Mastering Agile Tasks with Limited Trials
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- Dynamic Sampling that Adapts: Iterative DPO for Self-Aware Mathematical Reasoning
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models
- Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
- Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
- Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
- WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- Can Vision-Language Models Solve the Shell Game?
- Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
- Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos
- MinerU.Chem: A High-Precision System for Optical Chemical Structure and Reaction Recognition
- ChronoVision: Temporal Reasoning via Latent State Reconstruction
- MoCA: Implicit Social Context Analysis
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
Discussions
Related