vix.ing · top · new · best · stats · spec

Dhawan, Kunal

  1. Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
    2024/11/08 by Huang, Chien-yu, Chen, Wei-Chih, Yang, Shu-wen +77 · 25 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  2. Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
    2024/09/10 by Park, Taejin, Medennikov, Ivan, Dhawan, Kunal +6 · 12 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  3. Unified model for code-switching speech recognition and language identification based on a concatenated tokenizer
    2023/06/14 by Kunal Dhawan, Dhawan, Kunal, Dima Rekesh +3 · 4 citations
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Speech and dialogue systems
  4. Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach
    2023/09/11 by Tae Jin Park, Park, Tae Jin, Kunal Dhawan +5 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Topic Modeling #Music and Audio Processing
  5. Discrete Audio Representation as an Alternative to Mel-Spectrograms for Speaker and Speech Recognition
    2023/09/19 by Puvvada, Krishna C., Koluguri, Nithin Rao, Dhawan, Kunal +2 · 3 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  6. Training and Inference Efficiency of Encoder-Decoder Speech Models
    2025/03/07 by Żelasko, Piotr, Dhawan, Kunal, Galvez, Daniel +7 · 6 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
  7. Hindi-English Code-Switching Speech Corpus
    2018/09/24 by Ganji Sreeram, Kunal Dhawan, Sreeram, Ganji +3 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems
  8. Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs
    2024/06/07 by Ryan Langman, Langman, Ryan, Ante Jukić +6 · 2 citations
    Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. META-CAT: Speaker-Informed Speech Embeddings via Meta Information Concatenation for Multi-talker ASR
    2024/09/18 by Jinhan Wang, Weiqing Wang, Wang, Jinhan +17 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  10. NEST: Self-supervised Fast Conformer as All-purpose Seasoning to Speech Processing Tasks
    2024/08/23 by Huang, He, Park, Taejin, Dhawan, Kunal +6 · 2 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  11. Resource-Efficient Adaptation of Speech Foundation Models for Multi-Speaker ASR
    2024/09/02 by Wang, Weiqing, Dhawan, Kunal, Park, Taejin +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  12. SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
    2025/08/07 by Grossman, Raymond, Park, Taejin, Dhawan, Kunal +6 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Yang, Chao-Han Huck, Taejin Park +39 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  14. VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
    2024/10/23 by Yifan Peng, Peng, Yifan, Krishna C. Puvvada +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  15. Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
    2025/07/24 by Ivan Medennikov, Tae‐Jin Park, Medennikov, Ivan +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering