τ2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
2025/06/09 by Victor Barres, Honghua Dong, Barres, Victor +8 · 1 voice · 110 citations
Computer Science · Psychology · #AI in Service Interactions #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Social Robot Interaction and HRI #Speech and dialogue systems #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.2506.07982
openalex publication_date 2025/06/09 · arxiv published 2025/06/09 · arxiv updated 2025/06/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce τ2-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, τ2-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.
Citations
Cited by
- Close the Loop: Synthesizing Infinite Tool-Use Data via Multi-Agent Role-Playing
- GLM-5: from Vibe Coding to Agentic Engineering
- HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- PATHFinder Agent for Tailored Prenatal Care
- Cognitive Dark Matter: Measuring What AI Misses
- LLM Swiss Round: Aggregating Multi-Benchmark Performance via Competitive Swiss-System Dynamics
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
- Dynamic Tool Dependency Retrieval for Efficient Function Calling
- Towards a Science of Scaling Agent Systems
- Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction
- Qwen3.5-Omni Technical Report
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration
- Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
- AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models
- SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
- Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge
- MURMUR: Using cross-user chatter to break collaborative language agents in groups
- TPS-Bench: Evaluating AI Agents' Tool Planning & Scheduling Abilities in Compounding Tasks
- Simulating Environments with Reasoning Models for Agent Training
- The Ouroboros of Benchmarking: Reasoning Evaluation in an Era of Saturation
- ToolMind Technical Report: A Large-Scale, Reasoning-Enhanced Tool-Use Dataset
- LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
- SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
- Towards a Standard, Enterprise-Relevant Agentic AI Benchmark: Lessons from 5.5 billion tokens' worth of agentic AI evaluations
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviors
- TreeWriter: AI-Assisted Hierarchical Planning and Writing for Long-Form Documents
- Rethinking Text-to-SQL: Dynamic Multi-turn SQL Interaction for Real-world Database Exploration
- One Model to Critique Them All: Rewarding Agentic Tool-Use via Efficient Reasoning
- Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
- Experiential Reflective Learning for Self-Improving LLM Agents
- Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
- Living-Harness Is an Interactive-Agent Evolver
- SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
- CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
- OrchDAG: Complex Tool Orchestration in Multi-Turn Interactions with Plan DAGs
- APTBench: Benchmarking Agentic Potential of Base LLMs During Pre-Training
- BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
- A Survey on LLM Mid-Training
- On Generalization in Agentic Tool Calling: CoreThink Agentic Reasoner and MAVEN Dataset
- OlaMind: Towards Human-Like and Hallucination-Safe Customer Service for Retrieval-Augmented Dialogue
- A Coherence-Based Measure of AGI
- TheMCPCompany: Creating General-purpose Agents with Task-specific Tools
- KAT-Coder Technical Report
- AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- Empowering Real-World: A Survey on the Technology, Practice, and Evaluation of LLM-driven Industry Agents
- REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
- Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs
- ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
- COMPASS: Benchmarking Constrained Optimization in LLM Agents
- The Cognitive Bandwidth Bottleneck: Shifting Long-Horizon Agent from Planning with Actions to Planning with Schemas
- BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions
- Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
- Open Agent Specification (Agent Spec): A Unified Representation for AI Agents
- TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
- Non-Collaborative User Simulators for Tool Agents
- Who's Laughing Now? An Overview of Computational Humour Generation and Explanation
- UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
- LLMs struggle to simulate human belief updates in controlled environments
- VAmoS Bench: Voice Agent Simulation Bench
- PAUSE: A User-Centric Benchmark for Personal AI Assistants in Unified Service Environments
- Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks
- Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
- Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- Introducing LongCat-Flash-Thinking: A Technical Report
- Instruction-Following Evaluation in Function Calling for Large Language Models
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- LIMI: Less is More for Agency
- Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
- Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents
- Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
- Towards General Agentic Intelligence via Environment Scaling
- Reinforcement Learning Foundations for Deep Research Systems: A Survey
- LongCat-Flash Technical Report
- MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
- UserBench: An Interactive Gym Environment for User-Centric Agents
- Kimi K2: Open Agentic Intelligence
- AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
- PivotRL: High Accuracy Agentic Post-Training at Low Compute Cost
- Securing Agents With Tracked Capabilities
- Benchmarking LLM Tool-Use in the Wild
- Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments
- ERNIE 5.0 Technical Report
- Data Turnstile: A Scalable Open Framework for Function-Calling Data Generation
- Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
- DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
- When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary
- Progressive Agent Skill Generation via Reinforcement Learning
- The Real Barrier to LLM Agent Usability is Agentic ROI
- PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise
- Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce
- ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision
- IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations
- SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
- LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
- Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
- Cross-Benchmark Generalization in Long-Horizon Agents
- AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
- FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use
- When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
- Kwai Keye-VL-2.0 Technical Report
- ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
- Stateful Online Monitoring Catches Distributed Agent Attacks
- Scaling Embeddings Outperforms Scaling Experts in Language Models
Discussions
Related