vix.ing · top · new · best · stats · spec

Shi, Yangyang

  1. MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
    2024/02/22 by Zechun Liu, Liu, Zechun, Changsheng Zhao +21 · 3 voices · 42 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #cs.AI #cs.CL #cs.LG
  2. LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
    2023/05/29 by Zechun Liu, Barlas Oğuz, Liu, Zechun +15 · 32 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling
  3. Agent-as-a-Judge: Evaluate Agents with Agents
    2024/10/14 by Zhuge, Mingchen, Zhao, Changsheng, Ashley, Dylan +10 · 48 citations
    #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences
  4. Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion
    2021/04/05 by Duc Le, Le, Duc, Mahaveer Jain +21 · 11 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  5. Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
    2024/11/18 by Igor Fedorov, Fedorov, Igor, Kate Plawiak +37 · 2 voices · 5 citations
    #cs.DC #cs.AI
  6. TorchAudio: Building Blocks for Audio and Speech Processing
    2021/10/28 by Yao-Yuan Yang, Moto Hira, Yang, Yao-Yuan +43 · 8 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  7. FoleyGen: Visually-Guided Audio Generation
    2023/09/19 by Xinhao Mei, Varun Nagaraja, Mei, Xinhao +11 · 10 citations
    Computer Science · #Speech and Audio Processing #Music and Audio Processing #Generative Adversarial Networks and Image Synthesis
  8. Emformer: Efficient Memory Transformer Based Acoustic Model For Low Latency Streaming Speech Recognition
    2020/10/21 by Yangyang Shi, Yongqiang Wang, Shi, Yangyang +13 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  9. TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch
    2023/10/27 by Hwang, Jeff, Hira, Moto, Chen, Caroline +21 · 6 citations
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  10. High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
    2024/07/04 by Gaël Le Lan, Bowen Shi, Lan, Gael Le +21 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  11. Scaling Parameter-Constrained Language Models with Quality Data
    2024/10/04 by Chang, Ernie, Paltenghi, Matteo, Li, Yang +7 · 6 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  12. SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
    2024/12/03 by Haohe Liu, Liu, Haohe, Gaël Le Lan +17 · 6 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #electronic engineering #information engineering
  13. Transformer in action: a comparative study of transformer-based acoustic models for large scale speech recognition applications
    2020/10/27 by Wang, Yongqiang, Shi, Yangyang, Zhang, Frank +4 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences #Sound (cs.SD)
  14. Revisiting Sample Size Determination in Natural Language Understanding
    2023/07/01 by Ernie Chang, Chang, Ernie, Muhammad H. Rashid +11 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning and Algorithms #Natural Language Processing Techniques #Topic Modeling
  15. ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
    2025/02/04 by Zechun Liu, Changsheng Zhao, Liu, Zechun +28 · 5 citations
    Physics and Astronomy · Computer Science · #Particle Detector Development and Performance #Atomic and Subatomic Physics Research #Advanced Data Compression Techniques
  16. MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
    2025/01/15 by Junda Cheng, Cheng, Junda, Zhipeng Cai +18 · 5 citations
    Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Modular Robots and Swarm Intelligence
  17. Folding Attention: Memory and Power Optimization for On-Device Transformer-based Streaming Speech Recognition
    2023/09/14 by Yang Li, Liangzhen Lai, Li, Yang +12 · 2 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Hardware Architecture (cs.AR) #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
  18. Dissecting User-Perceived Latency of On-Device E2E Speech Recognition
    2021/04/06 by Shangguan, Yuan, Prabhavalkar, Rohit, Su, Hang +8 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  19. Streaming Transformer Transducer Based Speech Recognition Using Non-Causal Convolution
    2021/10/07 by Shi, Yangyang, Wu, Chunyang, Wang, Dilin +9 · 1 citation
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  20. LiCo-Net: Linearized Convolution Network for Hardware-efficient Keyword Spotting
    2022/11/09 by Yang, Haichuan, Yang, Zhaojun, Wan, Li +10 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #electronic engineering #information engineering
  21. Improving Fast-slow Encoder based Transducer with Streaming Deliberation
    2022/12/15 by Li, Ke, Mahadeokar, Jay, Guo, Jinxi +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #electronic engineering #information engineering
  22. Binary and Ternary Natural Language Generation
    2023/06/02 by Liu, Zechun, Oguz, Barlas, Pappu, Aasish +2 · 1 citation
    #Computation and Language (cs.CL) #FOS: Computer and information sciences
  23. DISGO: Automatic End-to-End Evaluation for Scene Text OCR
    2023/08/25 by Hwang, Mei-Yuh, Shi, Yangyang, Ramchandani, Ankit +6 · 1 citation
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  24. Enhance audio generation controllability through representation similarity regularization
    2023/09/15 by Shi, Yangyang, Lan, Gael Le, Nagaraja, Varun +6 · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering
  25. DepthLM: Metric Depth From Vision Language Models
    2025/09/29 by Zhipeng Cai, Ching-Feng Yeh, Cai, Zhipeng +17 · 4 citations
    Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques
  26. Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
    2024/06/13 by Frank Seide, Seide, Frank, Morrie Doulaty +9 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech and dialogue systems #electronic engineering #information engineering
  27. Breaking Down Power Barriers in On-Device Streaming ASR: Insights and Solutions
    2024/02/20 by Li, Yang, Shangguan, Yuan, Wang, Yuhao +5 · 1 citation
    #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #electronic engineering #information engineering