vix.ing · top · new · best · stats · spec

Philip C. Woodland

  1. Confidence Estimation for Attention-based Sequence-to-sequence Models for Speech Recognition
    2020/10/22 by Qiujia Li, Li, Qiujia, David Qiu +13 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  2. Can Contextual Biasing Remain Effective with Whisper and GPT-2?
    2023/06/02 by Guangzhi Sun, Xianrui Zheng, Sun, Guangzhi +5 · 5 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  3. CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
    2025/01/24 by Guangzhi Sun, 晓慧 战, Sun, Guangzhi +7 · 8 citations
    Computer Science · Social Sciences · #Access Control and Trust #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling
  4. Discrete Audio Tokens: More Than a Survey!
    2025/06/12 by Pooneh Mousavi, Gallil Maimon, Mousavi, Pooneh +39 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Digital Media Forensic Detection #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  5. Transformer Language Models with LSTM-based Cross-utterance Information Representation
    2021/02/12 by Guangzhi Sun, Sun, G., C. Zhang +3 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
  6. Adapting GPT, GPT-2 and BERT Language Models for Speech Recognition
    2021/07/29 by Xianrui Zheng, Chao Zhang, Zheng, Xianrui +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  7. Knowledge Distillation for Neural Transducers from Large Self-Supervised Pre-trained Models
    2021/10/07 by Xiaoyu Yang, Qiujia Li, Yang, Xiaoyu +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  8. 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem
    2024/05/30 by Mingjie Chen, Chen, Mingjie, H.L. Zhang +25 · 4 citations
    Engineering · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Vehicle License Plate Recognition #electronic engineering #information engineering
  9. CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
    2024/05/22 by Guangzhi Sun, Sun, Guangzhi, Potsawee Manakul +11 · 3 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topological and Geometric Data Analysis
  10. Decoupled Structure for Improved Adaptability of End-to-End Models
    2023/08/25 by Keqi Deng, Philip C. Woodland, Deng, Keqi +1 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  11. Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
    2024/06/06 by Keqi Deng, Philip C. Woodland, Deng, Keqi +1 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #electronic engineering #information engineering
  12. Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
    2022/05/18 by Guangzhi Sun, Chao Zhang, Sun, Guangzhi +3 · 1 citation
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  13. Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription
    2022/07/08 by Xianrui Zheng, Chao Zhang, Zheng, Xianrui +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  14. SkillAggregation: Reference-free LLM-Dependent Aggregation
    2024/10/14 by Guangzhi Sun, Sun, Guangzhi, Anmol Kagrecha +7 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #FOS: Computer and information sciences #Imbalanced Data Classification Techniques #Machine Learning (cs.LG) #Rough Sets and Fuzzy Logic
  15. Graph Neural Networks for Contextual ASR with the Tree-Constrained Pointer Generator
    2023/05/30 by Guangzhi Sun, Chao Zhang, Sun, Guangzhi +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  16. MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
    2024/09/25 by Xiaoyu Yang, Yang, Xiaoyu, Qiujia Li +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  17. SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
    2025/04/22 by Keqi Deng, Deng, Keqi, Wenxi Chen +5 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  18. Text-Prompted CLAP: Learning Query-Conditioned Audio Representations via Contrastive Learning
    2026/07/27 by Mohan Li, Rama Doddipatla, Philip C. Woodland
    #eess.AS