Kingsbury, Brian
- Everything at Once -- Multi-modal Fusion Transformer for Video Retrieval
2021/12/08 by Shvetsova, Nina, Chen, Brian, Rouditchenko, Andrew +6 · 7 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Beyond Backprop: Online Alternating Minimization with Auxiliary Variables
2018/06/24 by Anna Choromanska, Choromanska, Anna, Benjamin Cowen +19 · 6 citations
Computer Science · #Stochastic Gradient Optimization Techniques #Machine Learning and ELM #Domain Adaptation and Few-Shot Learning
- Estimating Information Flow in Deep Neural Networks
2018/10/12 by Ziv Goldfeld, E. van den Berg, Goldfeld, Ziv +11 · 3 citations
Computer Science · #Adversarial Robustness in Machine Learning #Stochastic Gradient Optimization Techniques #Domain Adaptation and Few-Shot Learning
- Comparison of Multilingual Self-Supervised and Weakly-Supervised Speech Pre-Training for Adaptation to Unseen Languages
2023/05/21 by Rouditchenko, Andrew, Khurana, Sameer, Thomas, Samuel +6 · 3 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- AVLnet: Learning Audio-Visual Language Representations from Instructional Videos
2020/06/16 by Rouditchenko, Andrew, Boggust, Angie, Harwath, David +11 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Sound (cs.SD) #electronic engineering #information engineering
- Advancing RNN Transducer Technology for Speech Recognition
2021/03/17 by Saon, George, Tueske, Zoltan, Bolanos, Daniel +1 · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- On the limit of English conversational speech recognition
2021/05/03 by Tüske, Zoltán, Saon, George, Kingsbury, Brian · 2 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Understanding Unequal Gender Classification Accuracy from Face Images
2018/11/30 by Vidya Muthukumar, Tejaswini Pedapati, Muthukumar, Vidya +17 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Face and Expression Recognition #Face recognition and analysis #Machine Learning (stat.ML)
- Towards Reducing the Need for Speech Training Data To Build Spoken Language Understanding Systems
2022/02/26 by Samuel Thomas, Hong-Kwang Jeff Kuo, Thomas, Samuel +5 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Semi-Autoregressive Streaming ASR With Label Context
2023/09/19 by Siddhant Arora, Arora, Siddhant, George Saon +5 · 3 citations
Computer Science · Engineering · #Advanced Chemical Sensor Technologies #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Energy Efficient Wireless Sensor Networks #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Improvements to deep convolutional neural networks for LVCSR
2013/09/05 by Sainath, Tara N., Kingsbury, Brian, Mohamed, Abdel-rahman +6 · 1 citation
#65K05 #90C15 #90C90 #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural and Evolutionary Computing (cs.NE) #Optimization and Control (math.OC)
- Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
2025/05/13 by George Saon, Saon, George, Avihu Dekel +44 · 7 citations
Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Semantic Web and Ontologies
- Challenging the Boundaries of Speech Recognition: The MALACH Corpus
2019/08/09 by Picheny, Michael, Tüske, Zóltan, Kingsbury, Brian +3 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- Single headed attention based sequence-to-sequence model for state-of-the-art results on Switchboard
2020/01/20 by Tüske, Zoltán, Saon, George, Audhkhasi, Kartik +1 · 1 citation
#68T10 #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #electronic engineering #information engineering
- Leveraging Unpaired Text Data for Training End-to-End Speech-to-Intent Systems
2020/10/08 by Huang, Yinghui, Kuo, Hong-Kwang, Thomas, Samuel +5 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- End-to-End Spoken Language Understanding Without Full Transcripts
2020/09/30 by Kuo, Hong-Kwang J., Tüske, Zoltán, Thomas, Samuel +7 · 1 citation
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #I.2.7 #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
- Integrating Dialog History into End-to-End Spoken Language Understanding\n Systems
2021/08/18 by Jatin Ganhotra, Samuel Thomas, Ganhotra, Jatin +11 · 1 citation
Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
- Joint Unsupervised and Supervised Training for Automatic Speech Recognition via Bilevel Optimization
2024/01/13 by Saif, A F M, Cui, Xiaodong, Shen, Han +3 · 3 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML)
- CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
2025/05/02 by Edson Araujo, Araujo, Edson, Andrew Rouditchenko +17 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Image and Signal Denoising Methods #Multimedia (cs.MM) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- High-Dimensional Smoothed Entropy Estimation via Dimensionality Reduction
2023/05/08 by Kristjan Greenewald, Greenewald, Kristjan, Brian Kingsbury +3 · 1 citation
Computer Science · Physics and Astronomy · #Adversarial Robustness in Machine Learning #FOS: Computer and information sciences #Information Theory (cs.IT) #Machine Learning (cs.LG) #Model Reduction and Neural Networks #Neural Networks and Applications
- Very Deep Multilingual Convolutional Neural Networks for LVCSR
2015/09/29 by Tom Sercu, Sercu, Tom, Christian Puhrsch +5 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing
- Soft Random Sampling: A Theoretical and Empirical Analysis
2023/11/21 by Xiaodong Cui, Cui, Xiaodong, Ashish Mittal +9 · 1 citation
Computer Science · Engineering · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Machine Learning and ELM #Sparse and Compressive Sensing Techniques
- A Non-autoregressive Model for Joint STT and TTS
2025/01/15 by Vishal Sunder, Brian Kingsbury, Sunder, Vishal +11 · 1 citation
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering