vix.ing · top · new · best · stats · spec

Shinji Watanabe

  1. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
    2023/04/25 by Rongjie Huang, Mingze Li, Huang, Rongjie +23 · 1 voice · 46 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling
  2. SUPERB: Speech processing Universal PERformance Benchmark
    2021/05/03 by Shu-Wen Yang, Po-Han Chi, Yang, Shu-wen +37 · 68 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  3. ESPnet: End-to-End Speech Processing Toolkit
    2018/03/30 by Shinji Watanabe, Watanabe, Shinji, Takaaki Hori +21 · 53 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  4. Deep clustering: Discriminative embeddings for segmentation and separation
    2015/08/18 by John R. Hershey, Hershey, John R., Zhuo Chen +5 · 26 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  5. Conditional Diffusion Probabilistic Model for Speech Enhancement
    2022/02/10 by Yen‐Ju Lu, Lu, Yen-Ju, Zhong-Qiu Wang +9 · 27 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  6. CHiME-6 Challenge:Tackling Multispeaker Speech Recognition for Unsegmented Recordings
    2020/04/20 by Shinji Watanabe, Watanabe, Shinji, Michael Mandel +39 · 22 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  7. TF-GridNet: Integrating Full- and Sub-Band Modeling for Speech Separation
    2022/11/22 by Zhong-Qiu Wang, Samuele Cornell, Wang, Zhong-Qiu +9 · 23 citations
    Computer Science · Engineering · #Speech and Audio Processing #Speech Recognition and Synthesis #Advanced Adaptive Filtering Techniques
  8. EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
    2024/06/10 by Julius Richter, Richter, Julius, Yi-Chiao Wu +13 · 36 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. TF-GridNet: Making Time-Frequency Domain Models Great Again for Monaural Speaker Separation
    2022/09/08 by Zhong-Qiu Wang, Wang, Zhong-Qiu, Samuele Cornell +9 · 21 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  10. End-to-End Neural Speaker Diarization with Permutation-Free Objectives
    2019/09/12 by Yusuke Fujita, Naoyuki Kanda, Fujita, Yusuke +7 · 14 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  11. A review of speaker diarization: Recent advances with deep learning
    2021/11/13 by Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis +3 · 17 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  12. Music ControlNet: Multiple Time-varying Controls for Music Generation
    2023/11/13 by Shih-Lun Wu, Wu, Shih-Lun, Chris Donahue +5 · 23 citations
    Computer Science · Neuroscience · #Music and Audio Processing #Music Technology and Sound Studies #Neuroscience and Music Perception
  13. HEAR: Holistic Evaluation of Audio Representations
    2022/03/06 by Joseph Turian, Jordie Shier, Turian, Joseph +43 · 14 citations
    Computer Science · Neuroscience · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  14. Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
    2023/09/18 by Chien‐Yu Huang, Ke-Han Lu, Huang, Chien-yu +27 · 19 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  15. E-Branchformer: Branchformer with Enhanced merging for speech recognition
    2022/09/30 by Kwangyoun Kim, Felix F. Wu, Kim, Kwangyoun +11 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  16. Efficient Sequence Transduction by Jointly Predicting Tokens and Durations
    2023/04/13 by Xu H, Xu, Hainan, Fei Jia +9 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
  17. SPGISpeech: 5,000 hours of transcribed financial audio for fully\n formatted end-to-end speech recognition
    2021/04/05 by Patrick O’Neill, Vitaly Lavrukhin, O'Neill, Patrick K. +24 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  18. ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
    2019/10/24 by Tomoki Hayashi, Hayashi, Tomoki, Ryuichi Yamamoto +15 · 9 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Music and Audio Processing
  19. End-to-End Neural Speaker Diarization with Self-attention
    2019/09/13 by Yusuke Fujita, Naoyuki Kanda, Fujita, Yusuke +9 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  20. SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
    2024/01/30 by Takaaki Saeki, Saeki, Takaaki, Soumi Maiti +7 · 16 citations
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  21. SpoofCeleb: Speech Deepfake Detection and SASV In The Wild
    2024/09/18 by Jee-weon Jung, Jung, Jee-weon, Yihan Wu +25 · 20 citations
    Computer Science · #Speech Recognition and Synthesis
  22. On The Landscape of Spoken Language Models: A Comprehensive Survey
    2025/04/11 by Siddhant Arora, Arora, Siddhant, Kai‐Wei Chang +17 · 34 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  23. TorchAudio: Building Blocks for Audio and Speech Processing
    2021/10/28 by Yao-Yuan Yang, Yang, Yao-Yuan, Moto Hira +43 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  24. Towards Robust Speech Representation Learning for Thousands of Languages
    2024/06/30 by William Chen, Wangyou Zhang, Chen, William +17 · 15 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
  25. SpeechLMScore: Evaluating speech generation using speech language model
    2022/12/08 by Soumi Maiti, Maiti, Soumi, Yifan Peng +5 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  26. The third ‘CHiME’ speech separation and recognition challenge: Analysis and outcomes
    2016/12/06 by Jon Barker, Ricard Marxer, Emmanuel Vincent +1 · 6 citations
    Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  27. YODAS: Youtube-Oriented Dataset for Audio and Speech
    2024/06/02 by Xinjian Li, Shinnosuke Takamichi, Li, Xinjian +9 · 13 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #Video Analysis and Summarization #electronic engineering #information engineering
  28. Self-Supervised Speech Representations are More Phonetic than Semantic
    2024/06/12 by Kwanghee Choi, Choi, Kwanghee, Ankita Pasad +9 · 12 citations
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis #Natural Language Processing Techniques
  29. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict
    2020/05/18 by Yosuke Higuchi, Higuchi, Yosuke, Shinji Watanabe +7 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  30. UnitY: Two-pass Direct Speech-to-speech Translation with Discrete Units
    2022/12/15 by Hirofumi Inaguma, Sravya Popuri, Inaguma, Hirofumi +17 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  31. Reproducing Whisper-Style Training Using an Open-Source Toolkit and Publicly Available Data
    2023/09/25 by Yifan Peng, Jinchuan Tian, Peng, Yifan +29 · 8 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  32. Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
    2025/03/03 by Siddhant Arora, Arora, Siddhant, Zhiyun Lu +7 · 1 voice · 12 citations
    Psychology · Computer Science · #Phonetics and Phonology Research #Music Technology and Sound Studies #Music and Audio Processing
  33. ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
    2024/01/30 by Jee-weon Jung, Jung, Jee-weon, Wangyou Zhang +13 · 9 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  34. Muskits: an End-to-End Music Processing Toolkit for Singing Voice Synthesis
    2022/05/09 by Jiatong Shi, Shi, Jiatong, Shuai Guo +21 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  35. Improving ASR Contextual Biasing with Guided Attention
    2024/01/16 by Jiyang Tang, Kwangyoun Kim, Tang, Jiyang +9 · 7 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  36. Advances in Joint CTC-Attention based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM
    2017/06/08 by Takaaki Hori, Hori, Takaaki, Shinji Watanabe +5 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
  37. Prompting the Hidden Talent of Web-Scale Speech Models for Zero-Shot Task Generalization
    2023/05/18 by Puyuan Peng, Brian Yan, Peng, Puyuan +5 · 5 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  38. Far-Field Automatic Speech Recognition
    2020/09/09 by Reinhold Haeb-Umbach, Reinhold Haeb‐Umbach, Jahn Heymann +4 · 5 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  39. End-to-End Multi-speaker Speech Recognition with Transformer
    2020/02/10 by Xuankai Chang, Chang, Xuankai, Wangyou Zhang +7 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  40. Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning
    2023/05/29 by Xuankai Chang, Brian Yan, Chang, Xuankai +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  41. Transformer ASR with Contextual Block Processing
    2019/10/16 by Emiru Tsunoo, Tsunoo, Emiru, Yosuke Kashiwagi +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  42. Investigating self-supervised learning for speech enhancement and separation
    2022/03/15 by Shinji Watanabe, Huang, Zili, Watanabe, Shinji +6 · 3 citations
    Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  43. Orthros: Non-autoregressive End-to-end Speech Translation with Dual-decoder
    2020/10/25 by Hirofumi Inaguma, Inaguma, Hirofumi, Yosuke Higuchi +7 · 4 citations
    Computer Science · #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  44. Discrete Audio Tokens: More Than a Survey!
    2025/06/12 by Pooneh Mousavi, Mousavi, Pooneh, Gallil Maimon +39 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  45. OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
    2025/02/14 by William Chen, Chen, William, Jinchuan Tian +9 · 9 citations
    Computer Science · #Speech Recognition and Synthesis
  46. OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
    2025/05/31 by Yifan Peng, Shakeel Muhammad, Peng, Yifan +11 · 13 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  47. Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
    2019/06/26 by Naoyuki Kanda, Shota Horiguchi, Kanda, Naoyuki +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  48. Reducing Barriers to Self-Supervised Learning: HuBERT Pre-training with Academic Compute
    2023/06/11 by William Chen, Chen, William, Xuankai Chang +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  49. UNSSOR: Unsupervised Neural Speech Separation by Leveraging Over-determined Training Mixtures
    2023/05/31 by Zhong-Qiu Wang, Shinji Watanabe, Wang, Zhong-Qiu +1 · 5 citations
    Computer Science · Engineering · Neuroscience · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Machine Learning (cs.LG) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
  50. Boosting Unknown-number Speaker Separation with Transformer Decoder-based Attractor
    2024/01/23 by Younglo Lee, Shukjae Choi, Lee, Younglo +7 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  51. DiscreTalk: Text-to-Speech as a Machine Translation Problem
    2020/05/12 by Tomoki Hayashi, Shinji Watanabe, Hayashi, Tomoki +1 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  52. The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
    2024/06/11 by Xuankai Chang, Jiatong Shi, Chang, Xuankai +17 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  53. Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
    2023/09/29 by Shih-Lun Wu, Wu, Shih-Lun, Xuankai Chang +11 · 3 citations
    Computer Science · Arts and Humanities · #Music and Audio Processing #Subtitles and Audiovisual Media #Speech Recognition and Synthesis
  54. Crystalline Electronic Field in Rare-Earth Based Quasicrystal and Approximant: Analysis of Quantum Critical Au–Al–Yb Quasicrystal and Approximant
    2021/04/27 by Shinji Watanabe, Mina Kawamoto · 4 citations
    Materials Science · #Quasicrystal Structures and Properties #Nanocluster Synthesis and Applications #X-ray Diffraction in Crystallography
  55. Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
    2024/02/01 by Zakaria Aldeneh, Takuya Higuchi, Aldeneh, Zakaria +15 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  56. Findings of the IWSLT 2024 Evaluation Campaign
    2024/11/07 by Ibrahim Said Ahmad, Ahmad, Ibrahim Said, Antonios Anastasopoulos +87 · 6 citations
    Computer Science · Engineering · #Advanced Data Processing Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #GNSS positioning and interference #Seismology and Earthquake Studies
  57. Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
    2025/02/10 by Kwanghee Choi, Eunjung Yeo, Choi, Kwanghee +8 · 1 voice · 5 citations
    Computer Science · Psychology · #Speech Recognition and Synthesis #Phonetics and Phonology Research
  58. Nonmagnetic Insulating States near the Mott Transitions on Lattices with Geometrical Frustration and Implications for κ-(ET)2Cu2(CN)3
    2002/03/01 by Hidekazu Morita, Shinji Watanabe, Masatoshi Imada · 1 citation
    Physics and Astronomy · #cond-mat.str-el #cond-mat.mtrl-sci
  59. ESPnet2-TTS: Extending the Edge of TTS Research
    2021/10/15 by Tomoki Hayashi, Hayashi, Tomoki, Ryuichi Yamamoto +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  60. A Study of Transducer based End-to-End ASR with ESPnet: Architecture, Auxiliary Loss and Decoding Strategies
    2022/01/14 by Florian Boyer, Yusuke Shinohara, Boyer, Florian +7 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Neural Networks and Applications #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  61. Aligning Text-to-Music Evaluation with Human Preferences
    2025/03/20 by Yichen Huang, Huang, Yichen, Zachary Novack +13 · 7 citations
    Computer Science · #Music and Audio Processing #Music Technology and Sound Studies #Speech and Audio Processing
  62. Contextualized Automatic Speech Recognition with Dynamic Vocabulary
    2024/05/22 by Yui Sudo, Sudo, Yui, Y. Fukumoto +8 · 1 voice · 2 citations
    Computer Science · #Speech Recognition and Synthesis
  63. Far-Field Automatic Speech Recognition
    2020/09/20 by Reinhold Haeb‐Umbach, Haeb-Umbach, Reinhold, Jahn Heymann +9 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  64. Less Peaky and More Accurate CTC Forced Alignment by Label Priors
    2024/04/22 by Ruizhe Huang, Huang, Ruizhe, Xiaohui Zhang +21 · 3 citations
    Computer Science · Engineering · #Advanced Numerical Analysis Techniques #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Handwritten Text Recognition Techniques #Machine Learning (cs.LG) #electronic engineering #information engineering
  65. Toward Universal Speech Enhancement for Diverse Input Conditions
    2023/09/29 by Wangyou Zhang, Zhang, Wangyou, Kohei Saijo +7 · 3 citations
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  66. Deep Speech Synthesis from Articulatory Representations
    2022/09/13 by Peter Wu, Wu, Peter, Shinji Watanabe +7 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Biological sciences #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Quantitative Methods (q-bio.QM) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  67. Speaker-Independent Acoustic-to-Articulatory Speech Inversion
    2023/02/14 by Peter Wu, Wu, Peter, Liwei Chen +11 · 2 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  68. Speech collage: code-switched audio generation by collaging monolingual corpora
    2023/09/27 by Amir Hussein, Hussein, Amir, Dorsa Zeinali +15 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  69. On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
    2024/06/13 by Jinchuan Tian, Yifan Peng, Tian, Jinchuan +9 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques
  70. Language model integration based on memory control for sequence to sequence speech recognition
    2018/11/06 by Jaejin Cho, Shinji Watanabe, Cho, Jaejin +11 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  71. AugSumm: towards generalizable speech summarization using synthetic labels from large language model
    2024/01/10 by Jee-weon Jung, Roshan Sharma, Jung, Jee-weon +7 · 2 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
  72. Multilingual End-to-End Speech Translation
    2019/10/01 by Hirofumi Inaguma, Kevin Duh, Inaguma, Hirofumi +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  73. End-to-End Automatic Speech Recognition Integrated With CTC-Based Voice Activity Detection
    2020/02/03 by Takenori Yoshimura, Tomoki Hayashi, Yoshimura, Takenori +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  74. Differentiable Allophone Graphs for Language-Universal Speech Recognition
    2021/07/24 by Brian Yan, Yan, Brian, Siddharth Dalmia +7 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  75. The 2020 ESPnet update: new features, broadened applications, performance improvements, and future plans
    2020/12/23 by Shinji Watanabe, Watanabe, Shinji, Florian Boyer +27 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  76. Summarizing Speech: A Comprehensive Survey
    2025/04/10 by Fabian Retkowski, Retkowski, Fabian, Maike Züfle +11 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #Video Analysis and Summarization #electronic engineering #information engineering
  77. Understanding the Tradeoffs in Client-side Privacy for Downstream Speech Tasks
    2021/01/22 by Peter Wu, Paul Pu Liang, Wu, Peter +9 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Privacy-Preserving Technologies in Data #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  78. Self-Guided Curriculum Learning for Neural Machine Translation
    2021/05/10 by Lei Zhou, Liang Ding, Zhou, Lei +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  79. Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
    2021/07/20 by Tianzi Wang, Yuya Fujita, Wang, Tianzi +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  80. SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
    2024/12/07 by Pengcheng Guo, Xuankai Chang, Guo, Pengcheng +7 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems
  81. Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
    2024/06/19 by Chenda Li, Li, Chenda, Samuele Cornell +5 · 2 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
  82. SRU++: Pioneering Fast Recurrence with Attention for Speech Recognition
    2021/10/11 by Jing Pan, Tao Lei, Pan, Jing +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  83. Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
    2021/09/09 by Hirofumi Inaguma, Yosuke Higuchi, Inaguma, Hirofumi +7 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech Recognition and Synthesis
  84. MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
    2024/06/14 by Jiatong Shi, Shi, Jiatong, Xutai Ma +7 · 2 citations
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis
  85. Fast-MD: Fast Multi-Decoder End-to-End Speech Translation with Non-Autoregressive Hidden Intermediates
    2021/09/27 by Hirofumi Inaguma, Siddharth Dalmia, Inaguma, Hirofumi +5 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  86. S3PRL-VC: Open-source Voice Conversion Framework with Self-supervised Speech Representations
    2021/10/12 by Wen-Chin Huang, Shuwen Yang, Huang, Wen-Chin +9 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  87. Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization
    2021/11/29 by Brian Yan, Yan, Brian, Chunlei Zhang +15 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  88. OpusLM: A Family of Open Unified Speech Language Models
    2025/06/21 by Jinchuan Tian, Tian, Jinchuan, William Chen +21 · 6 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Natural Language Processing Techniques
  89. Improving Speech Enhancement through Fine-Grained Speech Characteristics
    2022/07/01 by Muqiao Yang, Joseph Konan, Yang, Muqiao +9 · 1 citation
    Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  90. Online Continual Learning of End-to-End Speech Recognition Models
    2022/07/11 by Muqiao Yang, Ian Lane, Yang, Muqiao +3 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  91. On the Evaluation of Speech Foundation Models for Spoken Language Understanding
    2024/06/14 by Siddhant Arora, Ankita Pasad, Arora, Siddhant +21 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  92. ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
    2025/03/11 by Siddhant Arora, Arora, Siddhant, Yifan Peng +21 · 1 voice · 3 citations
    Computer Science · #Speech and dialogue systems #Multi-Agent Systems and Negotiation #Natural Language Processing Techniques
  93. Streaming Joint Speech Recognition and Disfluency Detection
    2022/11/16 by Hayato Futami, Emiru Tsunoo, Futami, Hayato +11 · 1 citation
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  94. Bridging Speech and Textual Pre-trained Models with Unsupervised ASR
    2022/11/06 by Jiatong Shi, Chan-Jan Hsu, Shi, Jiatong +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  95. Towards Zero-Shot Code-Switched Speech Recognition
    2022/11/02 by Brian Yan, Matthew Wiesner, Yan, Brian +7 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Natural Language Processing Techniques
  96. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
    2025/05/31 by Siddhant Arora, Arora, Siddhant, Jinchuan Tian +13 · 5 citations
    Computer Science · #Speech and dialogue systems #Intelligent Tutoring Systems and Adaptive Learning #Topic Modeling
  97. ESPnet-SpeechLM: An Open Speech Language Model Toolkit
    2025/02/21 by Jinchuan Tian, Jiatong Shi, Tian, Jinchuan +29 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #electronic engineering #information engineering
  98. BECTRA: Transducer-based End-to-End ASR with BERT-Enhanced Encoder
    2022/11/02 by Yosuke Higuchi, Higuchi, Yosuke, Tetsuji Ogawa +5 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  99. TAPLoss: A Temporal Acoustic Parameter Loss for Speech Enhancement
    2023/02/16 by Yunyang Zeng, Zeng, Yunyang, Joseph Konan +13 · 1 citation
    Computer Science · Psychology · #Speech and Audio Processing #Speech Recognition and Synthesis #Phonetics and Phonology Research
  100. Multi-Channel Target Speaker Extraction with Refinement: The WavLab Submission to the Second Clarity Enhancement Challenge
    2023/02/15 by Samuele Cornell, Zhong-Qiu Wang, Cornell, Samuele +9 · 1 citation
    Computer Science · Neuroscience · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  101. Exploration on HuBERT with Multiple Resolutions
    2023/06/01 by Jiatong Shi, Shi, Jiatong, Yun Tang +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  102. EFFUSE: Efficient Self-Supervised Feature Fusion for E2E ASR in Low Resource and Multilingual Scenarios
    2023/10/05 by Tejes Srivastava, Jiatong Shi, Srivastava, Tejes +5 · 1 citation
    Computer Science · Engineering · #Advanced Chemical Sensor Technologies #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Fault Detection and Control Systems #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  103. ESPnet-SLU: Advancing Spoken Language Understanding through ESPnet
    2021/11/29 by Siddhant Arora, Arora, Siddhant, Siddharth Dalmia +23 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  104. Layer Pruning on Demand with Intermediate CTC
    2021/06/17 by Jaesong Lee, Jingu Kang, Lee, Jaesong +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  105. Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
    2024/06/05 by Yui Sudo, Muhammad Shakeel, Sudo, Yui +11 · 2 citations
    Engineering · Physics and Astronomy · #Advanced Radiotherapy Techniques #Advancements in Photolithography Techniques #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  106. SUPERB @ SLT 2022: Challenge on Generalization and Efficiency of Self-Supervised Speech Representation Learning
    2022/10/16 by Tzu-hsun Feng, Annie Dong, Feng, Tzu-hsun +25 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  107. Improving Design of Input Condition Invariant Speech Enhancement
    2024/01/25 by Wangyou Zhang, Jee-weon Jung, Zhang, Wangyou +5 · 1 citation
    Computer Science · Engineering · #Speech and Audio Processing #Advanced Adaptive Filtering Techniques #Speech Recognition and Synthesis
  108. An Investigation of End-to-End Multichannel Speech Recognition for\n Reverberant and Mismatch Conditions
    2019/04/18 by Aswin Shanmugam Subramanian, Subramanian, Aswin Shanmugam, Xiaofei Wang +9 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  109. Avoid Overthinking in Self-Supervised Models for Speech Recognition
    2022/11/01 by Dan Berrebbi, Brian Yan, Berrebbi, Dan +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  110. DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
    2024/06/13 by Suwon Shon, Shon, Suwon, Kwangyoun Kim +9 · 1 citation
    Computer Science · #Speech and dialogue systems #Speech Recognition and Synthesis #Natural Language Processing Techniques
  111. Beyond Silence: Bias Analysis through Loss and Asymmetric Approach in Audio Anti-Spoofing
    2024/06/25 by Hye-jin Shim, Shim, Hye-jin, Md Sahidullah +7 · 1 citation
    Computer Science · #Digital Media Forensic Detection #Music and Audio Processing
  112. Emergence of Antigenic Variants in Bovine H5N1 Influenza Viruses
    2025/05/01 by Kei Miyakawa, Makoto Ota, Kaori Sano +8 · 1 voice · 1 citation
    Medicine · Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · #Influenza Virus Research Studies #Animal Disease Management and Epidemiology #Virus-based gene therapy research
  113. Multi-Convformer: Extending Conformer with Multiple Convolution Kernels
    2024/07/04 by Prabhu, Darshan, Yifan Peng, Peng, Yifan +4 · 1 citation
    Computer Science · #Neural Networks and Applications #Image and Signal Denoising Methods
  114. Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
    2024/08/17 by Samuele Cornell, Jordan Darefsky, Cornell, Samuele +5 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  115. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  116. Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
    2023/10/09 by Jiatong Shi, William Chen, Shi, Jiatong +23 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  117. VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
    2024/10/23 by Yifan Peng, Krishna C. Puvvada, Peng, Yifan +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  118. The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
    2025/05/20 by Ming Gao, Gao, Ming, Shilong Wu +15 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  119. Predicted versatile topological nodal magnons in Tb-based icosahedral quasicrystal 1/1 approximants
    2025/02/16 by Rintaro Eto, Eto, Rintaro, Masahito Mochizuki +3 · 1 citation
    Materials Science · Physics and Astronomy · #Quasicrystal Structures and Properties #X-ray Diffraction in Crystallography #Historical Astronomy and Related Studies
  120. Continuous Speech Separation Using Speaker Inventory for Long Multi-talker Recording
    2020/12/17 by Cong Han, Han, Cong, Yi Luo +19 · 1 citation
    Computer Science · #Speech and Audio Processing #Speech Recognition and Synthesis #Music and Audio Processing
  121. OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
    2025/07/18 by Shikhar Bharadwaj, Bharadwaj, Shikhar, Samuele Cornell +11 · 2 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Speech Recognition and Synthesis
  122. On-device Streaming Discrete Speech Units
    2025/06/02 by Kwanghee Choi, Choi, Kwanghee, Masao Someki +5 · 2 citations
    Computer Science · Engineering · Social Sciences · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #IPv6, Mobility, Handover, Networks, Security #Machine Learning (cs.LG) #Multimedia Communication and Technology #Sound (cs.SD) #Wireless Communication Networks Research #electronic engineering #information engineering
  123. Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
    2026/07/28 by Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3
    Computer Science · #cs.CL #cs.SD
  124. Ground-State Phase Diagram of the Effective Model with Antiferromagnetic Nearest-Neighbor and Ferromagnetic Next-Nearest-Neighbor Interactions in 1/1 Periodic Approximants to Icosahedral Quasicrystals
    2026/07/27 by Shinji Watanabe
    Materials Science · Physics and Astronomy · Mathematics · #Quasicrystal Structures and Properties #Advanced Condensed Matter Physics #Analytic and geometric function theory