Hasegawa-Johnson, Mark
- AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
2019/05/14 by Kaizhi Qian, Qian, Kaizhi, Yang Zhang +7 · 17 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Semantic Image Inpainting with Deep Generative Models
2016/07/26 by Yeh, Raymond A., Chen, Chen, Lim, Teck Yian +3 · 9 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion
2024/03/21 by Hee Suk Yoon, Eunseop Yoon, Yoon, Hee Suk +9 · 18 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Dilated Recurrent Neural Networks
2017/10/05 by Shiyu Chang, Yang Zhang, Chang, Shiyu +17 · 9 citations
Computer Science · #Advanced Neural Network Applications #Neural Networks and Applications #Machine Learning and ELM
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers
2022/04/20 by Qian, Kaizhi, Zhang, Yang, Gao, Heting +5 · 9 citations
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Unsupervised Speech Decomposition via Triple Information Bottleneck
2020/04/23 by Qian, Kaizhi, Zhang, Yang, Chang, Shiyu +2 · 3 citations
#Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Fast and Efficient MMD-based Fair PCA via Optimization over Stiefel Manifold
2021/09/23 by Jung-Hyun Lee, Lee, Junghyun, Gwangsu Kim +7 · 2 citations
Psychology · #Evolutionary Psychology and Human Behavior
- Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
2024/09/07 by Junkai Wu, Xulin Fan, Wu, Junkai +11 · 5 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
2024/02/10 by Jialu Li, Mark Hasegawa‐Johnson, Li, Jialu +3 · 4 citations
Computer Science · Health Professions · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Infant Health and Development #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition
2018/02/07 by Yang, Xuesong, Audhkhasi, Kartik, Rosenberg, Andrew +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Equivariance Discovery by Learned Parameter-Sharing
2022/04/07 by Raymond A. Yeh, Yuan-Ting Hu, Yeh, Raymond A. +5 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Data Classification #Time Series Analysis and Forecasting
- Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition
2022/01/26 by Piotr Żelasko, Żelasko, Piotr, Siyuan Feng +13 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition
2022/03/29 by Ni, Junrui, Wang, Liming, Gao, Heting +4 · 1 citation
#Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #electronic engineering #information engineering
- ConfPO: Exploiting Policy Model Confidence for Critical Token Selection in Preference Optimization
2025/06/10 by Hee Suk Yoon, Eunseop Yoon, Yoon, Hee Suk +7 · 4 citations
Computer Science · #Machine Learning and Data Classification #Explainable Artificial Intelligence (XAI) #Recommender Systems and Techniques
- Aligning ASR Evaluation with Human and LLM Judgments: Intelligibility Metrics Using Phonetic, Semantic, and NLI Approaches
2025/06/19 by Phukon, Bornali, Zheng, Xiuwen, Hasegawa-Johnson, Mark · 3 citations
Medicine · Psychology · Computer Science · #Voice and Speech Disorders #Phonetics and Phonology Research #Speech Recognition and Synthesis
- Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
2023/09/13 by Jialu Li, Mark Hasegawa‐Johnson, Li, Jialu +3 · 1 citation
Social Sciences · Health Professions · Psychology · #Child Development and Digital Technology #Infant Health and Development #Language Development and Disorders
- Towards Unsupervised Speech Recognition Without Pronunciation Models
2024/06/12 by Ni, Junrui, Wang, Liming, Zhang, Yang +4 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering