vix.ing · top · new · best · stats

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

2026/05/13 by Tara Bogavelli, Gabrielle Gauthier Melançon, Katrina Stankiewicz +10 · 1 voice · 3 citations
Computer Science · #cs.AI #cs.CL #cs.LG #cs.SD

paper · pdf · doi:10.48550/arxiv.2605.13841

arxiv published 2026/05/13 · arxiv updated 2026/05/27

Abstract

Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses two core evaluation challenges: generating realistic simulated conversations, and measuring quality across the full scope of voice-specific failure modes. We present EVA-Bench, an end-to-end evaluation framework that addresses both. On the simulation side, EVA-Bench orchestrates bot-to-bot audio conversations over dynamic multi-turn dialogues, with automatic simulation validation that detects user simulator error and appropriately regenerates conversations before scoring. On the measurement side, EVA-Bench introduces two composite metrics: EVA-A (Accuracy), capturing task completion, faithfulness, and audio-level speech fidelity; and EVA-X (Experience), capturing conversation progression, spoken conciseness, and turn-taking timing. Both metrics apply to all major agent architectures, enabling direct cross-architecture comparison. EVA-Bench includes 213 scenarios across three enterprise domains, a controlled perturbation suite for accent and noise robustness, and pass@1, pass@k, passk measurements that distinguish peak from reliable capability. Across 12 systems spanning all three architectures, we find: (1) no system simultaneously exceeds 0.5 on both EVA-A pass@1 and EVA-X pass@1; (2) peak and reliable performance diverge substantially (median pass@k--passk gap of 0.44 on EVA-A); and (3) accent and noise perturbations expose substantial robustness gaps, with effects varying across architectures, systems, and metrics (mean Δ up to 0.314). We release the full framework, evaluation suite, and benchmark data under an open-source license.

Cited by

Discussions

Related