vix.ing · top · new · best · stats · spec

James Glass

  1. DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
    2023/09/07 by Yung-Sung Chuang, Yujia Xie, Chuang, Yung-Sung +9 · 3 voices · 65 citations
    Computer Science · #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling #cs.AI #cs.CL #cs.LG
  2. AST: Audio Spectrogram Transformer
    2021/04/05 by Yuan Gong, Yu-An Chung, Gong, Yuan +3 · 102 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
  3. Listen, Think, and Understand
    2023/05/18 by Yuan Gong, Gong, Yuan, Hongyin Luo +7 · 35 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  4. SSAST: Self-Supervised Audio Spectrogram Transformer
    2021/10/19 by Yuan Gong, Gong, Yuan, Cheng-I Lai +5 · 24 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  5. Predicting Factuality of Reporting and Bias of News Media Sources
    2018/10/02 by Ramy Baly, Baly, Ramy, Georgi Karadzhov +7 · 1 voice · 5 citations
    Computer Science · Mathematics · #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #cs.IR #cs.LG #stat.ML
  6. Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
    2024/07/09 by Yung-Sung Chuang, Chuang, Yung-Sung, Linlu Qiu +9 · 2 voices · 24 citations
    Psychology · #cs.CL #cs.AI #cs.LG
  7. Contrastive Audio-Visual Masked Autoencoder
    2022/10/02 by Yuan Gong, Gong, Yuan, Andrew Rouditchenko +11 · 17 citations
    Computer Science · #Music and Audio Processing #Speech and Audio Processing #Video Analysis and Summarization
  8. An Unsupervised Autoregressive Model for Speech Representation Learning
    2019/04/05 by Yu-An Chung, Chung, Yu-An, Wei-Ning Hsu +5 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. DinoSR: Self-Distillation and Online Clustering for Self-supervised Speech Representation Learning
    2023/05/17 by Alexander H. Liu, Heng-Jui Chang, Liu, Alexander H. +7 · 8 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Speech and dialogue systems
  10. Identifying and Controlling Important Neurons in Neural Machine Translation
    2018/11/03 by Anthony Bau, Bau, Anthony, Yonatan Belinkov +9 · 8 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  11. DiffCSE: Difference-based Contrastive Learning for Sentence Embeddings
    2022/04/21 by Yung-Sung Chuang, Rumen Dangovski, Chuang, Yung-Sung +17 · 6 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  12. Unsupervised Learning of Disentangled and Interpretable Representations from Sequential Data
    2017/09/22 by Wei-Ning Hsu, Hsu, Wei-Ning, Yu Zhang +3 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  13. Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces
    2018/05/18 by Yu-An Chung, Chung, Yu-An, Wei‐Hung Weng +5 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  14. Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
    2025/05/14 by Andrew Rouditchenko, Rouditchenko, Andrew, Saurabhchand Bhati +11 · 20 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  15. FAKTA: An Automatic End-to-End Fact Checking System
    2019/06/07 by Moin Nadeem, Nadeem, Moin, Wei Fang +7 · 1 voice · 1 citation
    Computer Science · Social Sciences · #Topic Modeling #Advanced Text Analysis Techniques #Misinformation and Its Impacts
  16. Spoken Moments: Learning Joint Audio-Visual Representations from Video\n Descriptions
    2021/05/10 by Mathew Monfort, Monfort, Mathew, SouYoung Jin +11 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Music and Audio Processing #Sound (cs.SD) #Video Analysis and Summarization #electronic engineering #information engineering
  17. Vector-Quantized Autoregressive Predictive Coding
    2020/05/17 by Yu-An Chung, Chung, Yu-An, Hao Tang +3 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  18. Meta CLIP 2: A Worldwide Scaling Recipe
    2025/07/29 by Yung-Sung Chuang, Li Yang, Chuang, Yung-Sung +27 · 15 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
  19. Negative Training for Neural Dialogue Response Generation
    2019/03/06 by Tianxing He, He, Tianxing, James Glass +1 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  20. Analyzing Phonetic and Graphemic Representations in End-to-End Automatic\n Speech Recognition
    2019/07/09 by Yonatan Belinkov, Belinkov, Yonatan, Ahmed Ali +3 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  21. PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition
    2021/06/10 by Cheng-I Lai, Lai, Cheng-I Jeff, Alexander H. Liu +16 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  22. Decoding on Graphs: Faithful and Sound Reasoning on Knowledge Graphs through Generation of Well-Formed Chains
    2024/10/24 by Kun Li, Tianhua Zhang, Li, Kun +9 · 4 citations
    Computer Science · #Logic, Reasoning, and Knowledge #Semantic Web and Ontologies #Rough Sets and Fuzzy Logic
  23. Highway Long Short-Term Memory RNNs for Distant Speech Recognition
    2015/10/30 by Yu Zhang, Guoguo Chen, Zhang, Yu +9 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
  24. Logic Against Bias: Textual Entailment Mitigates Stereotypical Sentence Reasoning
    2023/03/10 by Hongyin Luo, James Glass, Luo, Hongyin +1 · 2 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling
  25. Self-supervised Fine-tuning for Improved Content Representations by Speaker-invariant Clustering
    2023/05/18 by Heng-Jui Chang, Alexander H. Liu, Chang, Heng-Jui +3 · 2 citations
    Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Natural Language Processing Techniques
  26. Adaptive Query Rewriting: Aligning Rewriters through Marginal Probability of Conversational Answers
    2024/06/16 by Tianhua Zhang, Kun Li, Zhang, Tianhua +9 · 3 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling #Speech and dialogue systems
  27. Disentangling by Partitioning: A Representation Learning Framework for\n Multimodal Sensory Data
    2018/05/29 by Wei-Ning Hsu, James Glass, Hsu, Wei-Ning +1 · 2 citations
    Computer Science · #Speech and Audio Processing
  28. Convolutional Neural Networks and Language Embeddings for End-to-End Dialect Recognition
    2018/03/12 by Suwon Shon, Ahmed Ali, Shon, Suwon +3 · 1 citation
    Computer Science · #Speech Recognition and Synthesis #Natural Language Processing Techniques #Topic Modeling
  29. Detecting egregious responses in neural sequence-to-sequence models
    2018/09/11 by Tianxing He, He, Tianxing, James Glass +1 · 2 citations
    Biochemistry, Genetics and Molecular Biology · #Cell Image Analysis Techniques
  30. State-Space Large Audio Language Models
    2024/11/24 by Saurabhchand Bhati, Yuan Gong, Bhati, Saurabhchand +9 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #electronic engineering #information engineering
  31. An Empirical Study on Few-shot Knowledge Probing for Pretrained Language Models
    2021/09/06 by Tianxing He, Kyunghyun Cho, He, Tianxing +3 · 1 citation
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  32. Generate, Discriminate, Evolve: Enhancing Context Faithfulness via Fine-Grained Sentence-Level Self-Evolution
    2025/03/03 by Ke Li, Tianhua Zhang, Li, Kun +13 · 3 citations
    Social Sciences · #Socioeconomic Development in MENA
  33. Developing a Series of AI Challenges for the United States Department of the Air Force
    2022/07/14 by Vijay Gadepally, Gregory Angelides, Gadepally, Vijay +81 · 1 citation
    Business, Management and Accounting · Decision Sciences · #Artificial Intelligence (cs.AI) #Big Data and Business Intelligence #Computers and Society (cs.CY) #Data Quality and Management #FOS: Computer and information sciences #Scientific Computing and Data Management
  34. UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
    2025/03/02 by Alexander H. Liu, Liu, Alexander H., Sang-gil Lee +13 · 3 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  35. DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
    2024/10/31 by Heng-Jui Chang, Hongyu Gong, Chang, Heng-Jui +7 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
  36. USAD: Universal Speech and Audio Representation via Distillation
    2025/06/23 by Heng-Jui Chang, Chang, Heng-Jui, Saurabhchand Bhati +5 · 4 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  37. Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
    2024/01/16 by Alexander H. Liu, Liu, Alexander H., Sung-Lin Yeh +3 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  38. Generative Pre-Training for Speech with Autoregressive Predictive Coding
    2019/10/23 by Yu-An Chung, James Glass, Chung, Yu-An +1 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  39. Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
    2018/03/23 by Yu-An Chung, James Glass, Chung, Yu-An +1 · 1 citation
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling
  40. GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
    2024/10/08 by M. Jehanzeb Mirza, Mirza, M. Jehanzeb, Mengjie Zhao +27 · 2 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  41. A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
    2024/10/29 by Alexander H. Liu, Qirui Wang, Liu, Alexander H. +5 · 1 citation
    Computer Science · #Neural Networks and Applications
  42. mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
    2025/02/03 by Andrew Rouditchenko, Samuel Thomas, Rouditchenko, Andrew +7 · 1 citation
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering