James Glass
- DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
2023/09/07 by Yung-Sung Chuang, Yujia Xie, Chuang, Yung-Sung +9 · 3 voices · 65 citations
Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
- AST: Audio Spectrogram Transformer
2021/04/05 by Yuan Gong, Yu-An Chung, Gong, Yuan +3 · 102 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- Listen, Think, and Understand
2023/05/18 by Yuan Gong, Gong, Yuan, Hongyin Luo +7 · 35 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- SSAST: Self-Supervised Audio Spectrogram Transformer
2021/10/19 by Yuan Gong, Gong, Yuan, Cheng-I Lai +5 · 24 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Predicting Factuality of Reporting and Bias of News Media Sources
2018/10/02 by Ramy Baly, Baly, Ramy, Georgi Karadzhov +7 · 1 voice · 5 citations
Computer Science · Mathematics · #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.IR #cs.LG #stat.ML
- Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
2024/07/09 by Yung-Sung Chuang, Chuang, Yung-Sung, Linlu Qiu +9 · 2 voices · 24 citations
Psychology · #cs.CL #cs.AI #cs.LG
- Contrastive Audio-Visual Masked Autoencoder
2022/10/02 by Yuan Gong, Gong, Yuan, Andrew Rouditchenko +11 · 17 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Video Analysis and Summarization
- An Unsupervised Autoregressive Model for Speech Representation Learning
2019/04/05 by Yu-An Chung, Chung, Yu-An, Wei-Ning Hsu +5 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
2023/05/17 by Alexander H. Liu, Heng-Jui Chang, Liu, Alexander H. +7 · 8 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Speech and dialogue systems
- Identifying and Controlling Important Neurons in Neural Machine Translation
2018/11/03 by Anthony Bau, Bau, Anthony, Yonatan Belinkov +9 · 8 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings
2022/04/21 by Yung-Sung Chuang, Rumen Dangovski, Chuang, Yung-Sung +17 · 6 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
2017/09/22 by Wei-Ning Hsu, Hsu, Wei-Ning, Yu Zhang +3 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces
2018/05/18 by Yu-An Chung, Chung, Yu-An, Wei‐Hung Weng +5 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
2025/05/14 by Andrew Rouditchenko, Rouditchenko, Andrew, Saurabhchand Bhati +11 · 20 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- FAKTA: An Automatic End-to-End Fact Checking System
2019/06/07 by Moin Nadeem, Nadeem, Moin, Wei Fang +7 · 1 voice · 1 citation
Computer Science · Social Sciences · #Topic Modeling #Advanced Text Analysis Techniques #Misinformation and Its Impacts
- Spoken Moments: Learning Joint Audio-Visual Representations from Video\n Descriptions
2021/05/10 by Mathew Monfort, Monfort, Mathew, SouYoung Jin +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Video Analysis and Summarization #electronic engineering #information engineering
- Vector-Quantized Autoregressive Predictive Coding
2020/05/17 by Yu-An Chung, Chung, Yu-An, Hao Tang +3 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Meta CLIP 2: A Worldwide Scaling Recipe
2025/07/29 by Yung-Sung Chuang, Li Yang, Chuang, Yung-Sung +27 · 15 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
- Negative Training for Neural Dialogue Response Generation
2019/03/06 by Tianxing He, He, Tianxing, James Glass +1 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Analyzing Phonetic and Graphemic Representations in End-to-End Automatic\n Speech Recognition
2019/07/09 by Yonatan Belinkov, Belinkov, Yonatan, Ahmed Ali +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
2021/06/10 by Cheng-I Lai, Lai, Cheng-I Jeff, Alexander H. Liu +16 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains
2024/10/24 by Kun Li, Tianhua Zhang, Li, Kun +9 · 4 citations
Computer Science · #Logic, Reasoning, and Knowledge #Semantic Web and Ontologies #Rough Sets and Fuzzy Logic
- Highway Long Short-Term Memory RNNs for Distant Speech Recognition
2015/10/30 by Yu Zhang, Guoguo Chen, Zhang, Yu +9 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence Reasoning
2023/03/10 by Hongyin Luo, James Glass, Luo, Hongyin +1 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
- Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering
2023/05/18 by Heng-Jui Chang, Alexander H. Liu, Chang, Heng-Jui +3 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
- Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers
2024/06/16 by Tianhua Zhang, Kun Li, Zhang, Tianhua +9 · 3 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech and dialogue systems
- Disentangling by Partitioning: A Representation Learning Framework for\n Multimodal Sensory Data
2018/05/29 by Wei-Ning Hsu, James Glass, Hsu, Wei-Ning +1 · 2 citations
Computer Science · #Speech and Audio Processing
- Convolutional Neural Networks and Language Embeddings for End-to-End Dialect Recognition
2018/03/12 by Suwon Shon, Ahmed Ali, Shon, Suwon +3 · 1 citation
Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
- Detecting egregious responses in neural sequence-to-sequence models
2018/09/11 by Tianxing He, He, Tianxing, James Glass +1 · 2 citations
Biochemistry, Genetics and Molecular Biology · #Cell Image Analysis Techniques
- State-Space Large Audio Language Models
2024/11/24 by Saurabhchand Bhati, Yuan Gong, Bhati, Saurabhchand +9 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #electronic engineering #information engineering
- An Empirical Study on Few-shot Knowledge Probing for Pretrained Language Models
2021/09/06 by Tianxing He, Kyunghyun Cho, He, Tianxing +3 · 1 citation
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution
2025/03/03 by Ke Li, Tianhua Zhang, Li, Kun +13 · 3 citations
Social Sciences · #Socioeconomic Development in MENA
- Developing a Series of AI Challenges for the United States Department of the Air Force
2022/07/14 by Vijay Gadepally, Gregory Angelides, Gadepally, Vijay +81 · 1 citation
Business, Management and Accounting · Decision Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computers and Society (cs.CY) #Data Quality and Management #FOS: Computer and information sciences #Scientific Computing and Data Management
- UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
2025/03/02 by Alexander H. Liu, Liu, Alexander H., Sang-gil Lee +13 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
2024/10/31 by Heng-Jui Chang, Hongyu Gong, Chang, Heng-Jui +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- USAD: Universal Speech and Audio Representation via Distillation
2025/06/23 by Heng-Jui Chang, Chang, Heng-Jui, Saurabhchand Bhati +5 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
2024/01/16 by Alexander H. Liu, Liu, Alexander H., Sung-Lin Yeh +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Generative Pre-Training for Speech with Autoregressive Predictive Coding
2019/10/23 by Yu-An Chung, James Glass, Chung, Yu-An +1 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
2018/03/23 by Yu-An Chung, James Glass, Chung, Yu-An +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling
- GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
2024/10/08 by M. Jehanzeb Mirza, Mirza, M. Jehanzeb, Mengjie Zhao +27 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
2024/10/29 by Alexander H. Liu, Qirui Wang, Liu, Alexander H. +5 · 1 citation
Computer Science · #Neural Networks and Applications
- mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
2025/02/03 by Andrew Rouditchenko, Samuel Thomas, Rouditchenko, Andrew +7 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering