vix.ing · top · new · best · stats · spec

Piotr Żelasko

  1. Study of Pre-processing Defenses against Adversarial Attacks on\n State-of-the-art Speaker Recognition Systems
    2021/01/21 by Sonal Joshi, Joshi, Sonal, Jesús Villalba +7 · 3 citations
    Computer Science · Engineering · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Geophysical Methods and Applications
  2. Chain-of-Thought Prompting for Speech Translation
    2024/09/17 by Ke Hu, Hu, Ke, Zhehuai Chen +13 · 6 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques
  3. SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
    2025/05/21 by Ke Hu, Hu, Ke, Ehsan Hosseini-Asl +17 · 11 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
  4. BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
    2024/06/28 by Zhehuai Chen, He Huang, Chen, Zhehuai +13 · 5 citations
    Computer Science · #68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #I.2.7 #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. Adversarial Attacks and Defenses for Speech Recognition Systems
    2021/03/31 by Piotr Żelasko, Żelasko, Piotr, Sonal Joshi +11 · 2 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #Cryptography and Security (cs.CR) #FOS: Computer and information sciences #FOS: Electrical engineering #Hate Speech and Cyberbullying Detection #Sound (cs.SD) #electronic engineering #information engineering
  6. Predicting Multi-Codebook Vector Quantization Indexes for Knowledge Distillation
    2022/10/31 by Liyong Guo, Guo, Liyong, Xiaoyu Yang +20 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
  7. Punctuation Prediction in Spontaneous Conversations: Can We Mitigate ASR\n Errors with Retrofitted Word Embeddings?
    2020/04/13 by Łukasz Augustyniak, Augustyniak, Łukasz, Piotr Szymański +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  8. Joint prediction of truecasing and punctuation for conversational speech\n in low-resource scenarios
    2021/09/13 by Raghavendra Pappagari, Piotr Żelasko, Pappagari, Raghavendra +7 · 1 citation
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
  9. Representation Learning to Classify and Detect Adversarial Attacks\n against Speaker and Speech Recognition Systems
    2021/07/09 by Jesús Villalba, Villalba, Jesús, Sonal Joshi +5 · 1 citation
    Computer Science · #Adversarial Robustness in Machine Learning #Anomaly Detection Techniques and Applications #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #Speech Recognition and Synthesis #electronic engineering #information engineering
  10. Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition
    2022/01/26 by Piotr Żelasko, Siyuan Feng, Żelasko, Piotr +13 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  11. Fast and parallel decoding for transducer
    2022/10/31 by Wei Kang, Liyong Guo, Kang, Wei +14 · 1 citation
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #DNA and Biological Computing #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Network Packet Processing and Optimization #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  12. Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
    2025/09/17 by Monica Sekoyan, Nithin Rao Koluguri, Sekoyan, Monica +13 · 5 citations
    Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Phonetics and Phonology Research #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  13. NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
    2024/11/08 by Yen‐Ting Lin, Lin, Yen-Ting, Zhehuai Chen +23 · 2 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques
  14. EMMeTT: Efficient Multimodal Machine Translation Training
    2024/09/20 by Piotr Żelasko, Żelasko, Piotr, Zhehuai Chen +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
  15. Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
    2024/09/15 by Chao-Han Huck Yang, Taejin Park, Yang, Chao-Han Huck +39 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  16. VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
    2024/10/23 by Yifan Peng, Krishna C. Puvvada, Peng, Yifan +17 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering