Yangyang Shi
- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
2024/02/22 by Zechun Liu, Liu, Zechun, Changsheng Zhao +21 · 3 voices · 42 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
- Neural Computers
2026/04/07 by Mingchen Zhuge, Changsheng Zhao, Haozhe Liu +16 · 11 voices
#cs.LG #cs.AI
- LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
2023/05/29 by Zechun Liu, Barlas Oğuz, Liu, Zechun +15 · 32 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
- Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
2021/04/05 by Duc Le, Mahaveer Jain, Le, Duc +21 · 11 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
2024/11/18 by Igor Fedorov, Fedorov, Igor, Kate Plawiak +37 · 2 voices · 5 citations
#cs.DC #cs.AI
- TorchAudio: Building Blocks for Audio and Speech Processing
2021/10/28 by Yao-Yuan Yang, Moto Hira, Yang, Yao-Yuan +43 · 8 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- FoleyGen: Visually-Guided Audio Generation
2023/09/19 by Xinhao Mei, Mei, Xinhao, Varun Nagaraja +11 · 10 citations
Computer Science · #Speech and Audio Processing #Music and Audio Processing #Generative Adversarial Networks and Image Synthesis
- Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition
2020/10/21 by Yangyang Shi, Yongqiang Wang, Shi, Yangyang +13 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
2024/07/04 by Gaël Le Lan, Lan, Gael Le, Bowen Shi +21 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
2024/12/03 by Haohe Liu, Gaël Le Lan, Liu, Haohe +17 · 6 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
- Revisiting Sample Size Determination in Natural Language Understanding
2023/07/01 by Ernie Chang, Chang, Ernie, Muhammad H. Rashid +11 · 2 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning and Algorithms #Natural Language Processing Techniques #Topic Modeling
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
2025/02/04 by Zechun Liu, Liu, Zechun, Changsheng Zhao +28 · 5 citations
Physics and Astronomy · Computer Science · #Particle Detector Development and Performance #Atomic and Subatomic Physics Research #Advanced Data Compression Techniques
- Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition
2023/09/14 by Yang Li, Liangzhen Lai, Li, Yang +12 · 2 citations
Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hardware Architecture (cs.AR) #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
2024/06/13 by Frank Seide, Morrie Doulaty, Seide, Frank +9 · 2 citations
Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
- DepthLM: Metric Depth From Vision Language Models
2025/09/29 by Zhipeng Cai, Ching-Feng Yeh, Cai, Zhipeng +17 · 5 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques