GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
2025/08/08 by GLM-4. 5 Team, Team, 5 Team +357 · 21 voices · 133 citations
Computer Science · #Cognitive Computing and Networks #Distributed and Parallel Computing Systems #Rough Sets and Fuzzy Logic #cs.CL
paper · pdf · doi:10.48550/arxiv.2508.06471
openalex publication_date 2025/08/08 · openalex created_date 2025/10/15 · openalex updated_date 2026/07/28
Abstract
We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning, GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd overall among all evaluated models and 2nd on agentic benchmarks. We release both GLM-4.5 (355B parameters) and a compact version, GLM-4.5-Air (106B parameters), to advance research in reasoning and agentic AI systems. Code, models, and more information are available at https://github.com/zai-org/GLM-4.5.
Citations
Cited by
- TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
- D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios
- DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
- ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
- PlotTwist: A Creative Plot Generation Framework with Small Language Models
- DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning
- When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
- IoUPD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models
- MARS: Multi-hop Adaptive Retrieval and SPARQL Generation for KGQA
- Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
- MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
- A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
- A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
- MIRA: Multimodal Iterative Reasoning Agent for Image Editing
- Agentic Entropy-Balanced Policy Optimization
- Nested Browser-Use Learning for Agentic Information Seeking
- AutoForge: Automated Environment Synthesis for Agentic Reinforcement Learning
- Agentic Software Issue Resolution with Large Language Models: A Survey
- DiRL: An Efficient Post-Training Framework for Diffusion Language Models
- Scale Weight Decay and Train Better
- Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection
- daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization
- LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
- Streaming Video Instruction Tuning
- Step-DeepResearch Technical Report
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- MoE Pathfinder: Trajectory-driven Expert Pruning
- TimeSeries2Report prompting enables adaptive large language model management of lithium-ion batteries
- Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
- ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
- CNFinBench: A Benchmark for Safety and Compliance of Large Language Models in Finance
- EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce
- LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
- Nanbeige4-3B Technical Report: Exploring the Frontier of Small Language Models
- Turbo-Muon: Accelerating Orthogonality-Based Optimization with Pre-Conditioning
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
- CuES: A Curiosity-driven and Environment-grounded Synthesis Framework for Agentic RL
- Dion2: A Simple Method to Shrink Matrix in Muon
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- Beyond High-Entropy Exploration: Correctness-Aware Low-Entropy Segment-Based Advantage Shaping for Reasoning LLMs
- Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
- Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
- RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
- M3-Bench: Multi-Modal, Multi-Hop, Multi-Threaded Tool-Using MLLM Agent Benchmark
- Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding Tasks
- CriticSearch: Fine-Grained Credit Assignment for Search Agents via a Retrospective Critic
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
- Uncovering Strategic Egoism Behaviors in Large Language Models
- ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
- Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
- AlphaResearch: Accelerating New Algorithm Discovery with Language Models
- RedOne 2.0: Rethinking Domain-specific LLM Post-Training in Social Networking Services
- Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
- Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling
- Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilers
- MicroRemed: Benchmarking LLMs in Microservices Remediation
- MemSearcher: Training LLMs to Reason, Search and Manage Memory via End-to-End Reinforcement Learning
- Interaction as Intelligence Part II: Asynchronous Human-Agent Rollout for Long-Horizon Task Training
- MedVLSynther: Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
- MemTX: Transactional Belief Commit for Stateful Agent Memory
- Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?
- Tongyi DeepResearch Technical Report
- AgentFold: Long-Horizon Web Agents with Proactive Context Management
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
- On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
- Multi-turn Training with Basic Human Feedback Helps Little on LLM Reasoning
- BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills
- MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- MARS-M: When Variance Reduction Meets Matrices
- ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
- FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
- Automated Network Protocol Testing with LLM Agents
- A2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- A Survey on Agentic Multimodal Large Language Models
- PaperArena: An Evaluation Benchmark for Tool-Augmented Agentic Reasoning on Scientific Literature
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- Effective Strategies for Asynchronous Software Engineering Agents
- InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
- How Many Code and Test Cases Are Enough? Evaluating Test Cases Generation from a Binary-Matrix Perspective
- Beyond Turn Limits: Training Deep Search Agents with Dynamic Context Window
- Automating Android Build Repair: Bridging the Reasoning-Execution Gap in LLM Agents with Domain-Specific Tools
- The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
- Pushing Test-Time Scaling Limits of Deep Search with Asymmetric Verification
- A Set of Quebec-French Corpus of Regional Expressions and Terms
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
- Read Between the Lines: A Benchmark for Uncovering Political Bias in Bangla News Articles
- AudioToolAgent: An Agentic Framework for Audio-Language Models
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners
- InfoAgent: Advancing Autonomous Information-Seeking Agents
- SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression
- Towards a Comprehensive Scaling Law of Mixture-of-Experts
- Effective Quantization of Muon Optimizer States
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
- Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Talking Trees: Reasoning-Assisted Induction of Decision Trees for Tabular Data
- Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
- V-GameGym: Visual Game Generation for Code Large Language Models
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning
- SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation
- IFHierBench: Hierarchical Instruction Following for Large Language Models
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- Introducing LongCat-Flash-Thinking: A Technical Report
- LIMI: Less is More for Agency
- Single-stream Policy Optimization
- WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction
- Scaling Agents via Continual Pre-training
- Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- Baichuan-M2: Scaling Medical Capability with Large Verifier System
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- A.S.E: A Repository-Level Benchmark for Evaluating Security in AI-Generated Code
- Latent Inter-User Difference Modeling for LLM Personalization
- Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
- CodegenBench: Can LLMs Write Efficient Code Across Architectures?
Discussions
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] [hn, 417 points, 83 comments]
- GLM-4.5 paper arrived! Many people consider this the best Chinese model, I still favor K2, but I get it Another deepseek R1 descendant. Almost the same # of active params as K2 & R1 but with far fewer [bsky, 18 points, 2 comments]
- glm 4.5 paper has this arxiv.org/abs/2508.06471 [bsky, 10 points, 2 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models [hn, 3 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] [Discussion] [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models [pdf] [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models [pdf] #HackerNews https://www.arxiv.org/pdf/2508.06471 [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 https://news.ycombinator.com/item?id=44871337 [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 www.arxiv.org [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] View Article | Join the HN Conversation Summary of HN discussion 🧵👇 #hacker-news [bsky, 0 points, 1 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 [comments] [133 points] [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 (https://news.ycombinator.com/item?id=44871337) [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 (http://news.ycombinator.com/item?id=44871337) [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 (http://news.ycombinator.com/item?id=44871337) [bsky, 0 points, 0 comments]
- anybody tried z.ai in anger? not sure about the quality other than the claims; there is a more in-depth paper: https://arxiv.org/abs/2508.06471 but the pricing looks really hard to beat and you can ju [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models [pdf] view on hacker news [bsky, 0 points, 0 comments]
- 📰 GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] 🔗 https://www.arxiv.org/pdf/2508.06471 💬 Discuss on HN [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 (https://news.ycombinator.com/item?id=44871337) [bsky, 0 points, 0 comments]
- GLM-4.5: Agentic, Reasoning, and Coding (Arc) Foundation Models [pdf] https://www.arxiv.org/pdf/2508.06471 (https://news.ycombinator.com/item?id=44871337) [bsky, 0 points, 0 comments]
- https://bsky.app/profile/buzzing.cc.web.brid.gy/post/3lw6obckjza52 [bsky, 0 points, 0 comments]
Related