Mark Hasegawa‐Johnson
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
2019/05/14 by Kaizhi Qian, Qian, Kaizhi, Yang Zhang +7 · 32 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Dilated Recurrent Neural Networks
2017/10/05 by Shiyu Chang, Yang Zhang, Chang, Shiyu +17 · 20 citations
Computer Science · #Advanced Neural Network Applications #Neural Networks and Applications #Machine Learning and ELM
- Semantic Image Inpainting with Deep Generative Models
2016/07/26 by Raymond A. Yeh, Yeh, Raymond A., Chen Chen +9 · 15 citations
Computer Science · #Advanced Image Processing Techniques #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis
- C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
2024/03/21 by Hee Suk Yoon, Yoon, Hee Suk, Eunseop Yoon +9 · 29 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers
2022/04/20 by Kaizhi Qian, Qian, Kaizhi, Zhang, Yang +12 · 14 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Unsupervised Speech Decomposition via Triple Information Bottleneck
2020/04/23 by Kaizhi Qian, Qian, Kaizhi, Shiyu Chang +6 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SpeechSplit 2.0: Unsupervised speech disentanglement for voice conversion Without tuning autoencoder Bottlenecks
2022/03/26 by Chak Ho Chan, Kaizhi Qian, Chan, Chak Ho +4 · 5 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Fast and Efficient MMD-based Fair PCA via Optimization over Stiefel Manifold
2021/09/23 by Jung-Hyun Lee, Gwangsu Kim, Lee, Junghyun +7 · 4 citations
Psychology · #Evolutionary Psychology and Human Behavior
- TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
2024/07/23 by Eunseop Yoon, Hee Suk Yoon, Yoon, Eunseop +17 · 12 citations
Computer Science · Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Muscle activation and electromyography studies #Reinforcement Learning in Robotics
- Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
2023/09/13 by Jialu Li, Li, Jialu, Mark Hasegawa‐Johnson +3 · 5 citations
Social Sciences · Health Professions · Psychology · #Child Development and Digital Technology #Infant Health and Development #Language Development and Disorders
- Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
2024/02/10 by Jialu Li, Li, Jialu, Mark Hasegawa‐Johnson +3 · 5 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
2024/09/07 by Junkai Wu, Wu, Junkai, Xulin Fan +11 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Streaming Recommender Systems
2016/07/21 by Shiyu Chang, Chang, Shiyu, Jiliang Tang +10 · 1 citation
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Machine Learning and Algorithms #Recommender Systems and Techniques #Social and Information Networks (cs.SI)
- Equivariance Discovery by Learned Parameter-Sharing
2022/04/07 by Raymond A. Yeh, Yeh, Raymond A., Yuan-Ting Hu +5 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Time Series Analysis and Forecasting
- Deep Learning Based Speech Beamforming
2018/02/15 by Kaizhi Qian, Yang Zhang, Qian, Kaizhi +9 · 1 citation
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Indoor and Outdoor Localization Technologies #Signal Processing (eess.SP) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- When CTC Training Meets Acoustic Landmarks
2018/11/05 by Di He, He, Di, Xuesong Yang +9 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition
2022/01/26 by Piotr Żelasko, Żelasko, Piotr, Siyuan Feng +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Dual-Path Cross-Modal Attention for better Audio-Visual Speech Extraction
2022/07/09 by Zhongweiyang Xu, Xulin Fan, Xu, Zhongweiyang +3 · 1 citation
Computer Science · Engineering · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition
2022/03/29 by Junrui Ni, Liming Wang, Ni, Junrui +10 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
2025/06/10 by Hee Suk Yoon, Yoon, Hee Suk, Eunseop Yoon +7 · 4 citations
Computer Science · #Machine Learning and Data Classification #Explainable Artificial Intelligence (XAI) #Recommender Systems and Techniques
- LI-TTA: Language Informed Test-Time Adaptation for Automatic Speech Recognition
2024/08/11 by Eunseop Yoon, Yoon, Eunseop, Hee Suk Yoon +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Towards Unsupervised Speech Recognition Without Pronunciation Models
2024/06/12 by Junrui Ni, Liming Wang, Ni, Junrui +11 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering