Florian Metze
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
2021/09/28 by Hu Xu, Gargi Ghosh, Xu, Hu +13 · 81 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
- Masked Autoencoders that Listen
2022/07/13 by Po-Yao Huang, Xu Hu, Huang, Po-Yao +13 · 72 citations
Computer Science · #Music and Audio Processing #Speech and Audio Processing #Speech Recognition and Synthesis
- How2Sign: A Large-scale Multimodal Dataset for Continuous American Sign\n Language
2020/08/18 by Amanda Duarte, Shruti Palaskar, Duarte, Amanda +13 · 39 citations
Computer Science · Psychology · #Hand Gesture Recognition Systems #Hearing Impairment and Communication #Human Pose and Action Recognition
- Universal Phone Recognition with a Multilingual Allophone System
2020/02/26 by Xinjian Li, Li, Xinjian, Siddharth Dalmia +19 · 15 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- EESEN: End-to-End Speech Recognition using Deep RNN Models and WFST-based Decoding
2015/07/29 by Yajie Miao, Mohammad Gowayyed, Miao, Yajie +3 · 19 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing
- How2: A Large-scale Dataset for Multimodal Language Understanding
2018/11/01 by Ramon Sanabria, Ozan Çağlayan, Sanabria, Ramon +11 · 12 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
- Embodied AI Agents: Modeling the World
2025/06/27 by Pascale Fung, Yoram Bachrach, Fung, Pascale +39 · 1 voice · 21 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #cs.AI
- VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding
2021/05/20 by Hu Xu, Xu, Hu, Gargi Ghosh +13 · 7 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Human Pose and Action Recognition
- Support-set bottlenecks for video-text representation learning
2020/10/06 by Mandela Patrick, Po-Yao Huang, Patrick, Mandela +11 · 5 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
- A Comparison of Five Multiple Instance Learning Pooling Functions for Sound Event Detection with Weak Labeling
2018/10/22 by Yun Wang, Wang, Yun, Juncheng Li +3 · 3 citations
Computer Science · Environmental Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hydrological Forecasting Using AI #Music and Audio Processing #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering
- CTC Alignments Improve Autoregressive Translation
2022/10/11 by Brian Yan, Yan, Brian, Siddharth Dalmia +11 · 3 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- SANTLR: Speech Annotation Toolkit for Low Resource Languages
2019/08/02 by Xinjian Li, Li, Xinjian, Zhong Zhou +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #electronic engineering #information engineering
- Space-Time Crop & Attend: Improving Cross-modal Video Representation Learning
2021/03/18 by Mandela Patrick, Patrick, Mandela, Yuki M. Asano +11 · 2 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
- Searchable Hidden Intermediates for End-to-End Models of Decomposable Sequence Tasks
2021/05/02 by Siddharth Dalmia, Dalmia, Siddharth, Brian Yan +7 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Speech Summarization using Restricted Self-Attention
2021/10/12 by Roshan Sharma, Sharma, Roshan, Shruti Palaskar +5 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Topic Modeling #electronic engineering #information engineering
- Cross-Attention End-to-End ASR for Two-Party Conversations
2019/07/24 by Suyoun Kim, Kim, Suyoun, Siddharth Dalmia +3 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Dialog-context aware end-to-end speech recognition
2018/08/07 by Suyoun Kim, Florian Metze, Kim, Suyoun +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling
- Acoustic-to-Word Recognition with Sequence-to-Sequence Models
2018/07/23 by Shruti Palaskar, Palaskar, Shruti, Florian Metze +1 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Hierarchical Multi Task Learning With CTC
2018/07/18 by Ramon Sanabria, Sanabria, Ramon, Florian Metze +1 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Speech Recognition and Synthesis
- Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion
2019/06/27 by Suyoun Kim, Siddharth Dalmia, Kim, Suyoun +3 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and dialogue systems #Topic Modeling #electronic engineering #information engineering
- Differentiable Allophone Graphs for Language-Universal Speech Recognition
2021/07/24 by Brian Yan, Yan, Brian, Siddharth Dalmia +7 · 2 citations
Computer Science · #Speech Recognition and Synthesis #Music and Audio Processing #Speech and Audio Processing
- On Compositionality in Neural Machine Translation
2019/11/04 by Vikas Raunak, Raunak, Vikas, Vaibhav Kumar +3 · 2 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling #Multimodal Machine Learning Applications
- LegoNN: Building Modular Encoder-Decoder Models
2022/06/07 by Siddharth Dalmia, Dalmia, Siddharth, Dmytro Okhonko +12 · 1 citation
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering