vix.ing · top · new · best · stats · spec

Yuekai Zhang

  1. SPGISpeech: 5,000 hours of transcribed financial audio for fully\n formatted end-to-end speech recognition
    2021/04/05 by Patrick O’Neill, Vitaly Lavrukhin, O'Neill, Patrick K. +24 · 18 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  2. TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch
    2023/10/27 by Jeff Hwang, Hwang, Jeff, Moto Hira +45 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  3. Sequence-to-sequence Singing Voice Synthesis with Perceptual Entropy Loss
    2020/10/22 by Jiatong Shi, Shi, Jiatong, Shuai Guo +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. PINNsAgent: Automated PDE Surrogation with Large Language Models
    2025/01/21 by Qingpo Wuwu, Wuwu, Qingpo, Chonghan Gao +15 · 7 citations
    Computer Science · #Computational Engineering #FOS: Computer and information sciences #Finance #Intelligent Tutoring Systems and Adaptive Learning #Natural Language Processing Techniques #Topic Modeling #and Science (cs.CE)
  5. Tiny Transducer: A Highly-efficient Speech Recognition Model on Edge Devices
    2021/01/18 by Yuekai Zhang, Sining Sun, Zhang, Yuekai +3 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  6. TrimTail: Low-Latency Streaming ASR with Simple but Effective Spectrogram-Level Length Penalty
    2022/11/01 by Xingchen Song, Di Wu, Song, Xingchen +15 · 2 citations
    Engineering · Social Sciences · #Advanced Chemical Sensor Technologies #Advanced Computing and Algorithms #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Sound (cs.SD) #Underwater Vehicles and Communication Systems #electronic engineering #information engineering
  7. TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
    2024/12/11 by Xingchen Song, Song, Xingchen, Mengtao Xing +20 · 4 citations
    Computer Science · Decision Sciences · #Context-Aware Activity Recognition Systems #Personal Information Management and User Behavior
  8. ESPnet-SLU: Advancing Spoken Language Understanding through ESPnet
    2021/11/29 by Siddhant Arora, Siddharth Dalmia, Arora, Siddhant +23 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering