ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
2025/05/30 by Jiatong Shi, Yifan Cheng, Shi, Jiatong +15
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2505.24518
openalex publication_date 2025/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual Evaluation of Speech Quality), STOI (Short-Time Objective Intelligibility), and MOS (Mean Opinion Score) each capture different aspects of speech quality. However, these metrics often have different scales, assumptions, and dependencies, making joint estimation non-trivial. To address these issues, we introduce ARECHO (Autoregressive Evaluation via Chain-based Hypothesis Optimization), a chain-based, versatile evaluation system for speech assessment grounded in autoregressive dependency modeling. ARECHO is distinguished by three key innovations: (1) a comprehensive speech information tokenization pipeline; (2) a dynamic classifier chain that explicitly captures inter-metric dependencies; and (3) a two-step confidence-oriented decoding algorithm that enhances inference reliability. Experiments demonstrate that ARECHO significantly outperforms the baseline framework across diverse evaluation scenarios, including enhanced speech analysis, speech generation evaluation, and, noisy speech evaluation. Furthermore, its dynamic dependency modeling improves interpretability by capturing inter-metric relationships. Across tasks, ARECHO offers reference-free evaluation using its dynamic classifier chain to support subset queries (single or multiple metrics) and reduces error propagation via confidence-oriented decoding.
Citations
- Uni-VERSA: Versatile Speech Assessment with a Unified Network
- SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning
- QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
- Mellow: a small audio language model for reasoning
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
- Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
- Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
- VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
- Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
- MOS-Bench: Benchmarking Generalization Abilities of Subjective Speech Quality Assessment Models
- Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
- SCOREQ: Speech Quality Assessment with Contrastive Regression
- ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
- What Are They Doing? Joint Audio-Speech Co-Reasoning
- The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
- The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
- GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
- SingMOS: An extensive Open-Source Singing Voice Dataset for MOS Prediction
- Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
- Rotary Position Embedding for Vision Transformer
- EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
- PAM: Prompting Audio-Language Models for Audio Quality Assessment
- ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
- SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
- emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
- Multi-CMGAN+/+: Leveraging Multi-Objective Speech Quality Metric Prediction for Speech Enhancement
- NOMAD: Unsupervised Learning of Perceptual Embeddings for Speech Enhancement and Non-matching Reference Audio Quality Assessment
- Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
- Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
- PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions
- Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
- PLCMOS -- a data-driven non-intrusive metric for the evaluation of packet loss concealment algorithms
- TorchAudio-Squim: Reference-less Speech Quality and Intelligibility measures in TorchAudio
- ICASSP 2023 Deep Noise Suppression Challenge
- InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt
- Robust Speech Recognition via Large-Scale Weak Supervision
- MTI-Net: A Multi-Target Speech Intelligibility Prediction Model
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
- ConferencingSpeech 2022 Challenge: Non-intrusive Objective Speech Quality Assessment (NISQA) Challenge for Online Conferencing Applications
- The VoiceMOS Challenge 2022
- InQSS: a speech intelligibility and quality assessment model using a multi-task learning network
- Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
- WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing
- ESPnet2-TTS: Extending the Edge of TTS Research
- DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors
- AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks
- NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
- SUPERB: Speech processing Universal PERformance Benchmark
- Convolutive Transfer Function Invariant SDR training criteria for\n Multi-Channel Reverberant Speech Separation
- DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality metric to\n evaluate Noise Suppressors
- ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric
- Common Voice: A Massively-Multilingual Speech Corpus
- WHAM!: Extending Speech Separation to Noisy Environments
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
- Rank consistent ordinal regression for neural networks with application to age estimation
- AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline
- Order-Free RNN with Visual Attention for Multi-Label Classification
- Subjective comparison and evaluation of speech enhancement algorithms
- Image method for efficiently simulating small-room acoustics
Related