Yinmin Zhong
- DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference
2026/02/25 by Yongtong Wu, Shaoyuan Chen, Yinmin Zhong +10 · 6 voices
Computer Science · #cs.DC
- MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
2024/02/23 by Ziheng Jiang, Haibin Lin, Jiang, Ziheng +61 · 29 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
- Fast Distributed Inference Serving for Large Language Models
2023/05/10 by Bingyang Wu, Yinmin Zhong, Wu, Bingyang +9 · 18 citations
Computer Science · Engineering · #Advanced Graph Neural Networks #Distributed #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
- LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
2024/04/15 by Bingyang Wu, Shengyu Liu, Wu, Bingyang +9 · 24 citations
Computer Science · #Advanced Neural Network Applications #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Parallel Computing and Optimization Techniques #Topic Modeling #and Cluster Computing (cs.DC)
- Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
2025/02/17 by Ailin Huang, Huang, Ailin, Boyong Wu +299 · 1 voice · 35 citations
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Sound (cs.SD) #Speech and dialogue systems #cs.AI #cs.CL #cs.HC #cs.SD #eess.AS #electronic engineering #information engineering
- Optimizing RLHF Training for Large Language Models with Stage Fusion
2024/09/20 by Yinmin Zhong, Zhong, Yinmin, Zili Zhang +18 · 18 citations
Computer Science · #Computation and Language (cs.CL) #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #Speech Recognition and Synthesis #Topic Modeling #and Cluster Computing (cs.DC)
- FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
2024/06/11 by Li‐Wen Chang, Wenlei Bao, Chang, Li-Wen +22 · 13 citations
Computer Science · Neuroscience · #Parallel Computing and Optimization Techniques #Embedded Systems Design Techniques #Brain Tumor Detection and Classification
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
2023/02/22 by Zhuohan Li, Lianmin Zheng, Li, Zhuohan +19 · 7 citations
Computer Science · #Age of Information Optimization #Cloud Computing and Resource Management #Distributed #FOS: Computer and information sciences #IoT and Edge/Fog Computing #Machine Learning (cs.LG) #Networking and Internet Architecture (cs.NI) #Parallel #and Cluster Computing (cs.DC)
- StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
2025/04/22 by Yinmin Zhong, Zhong, Yinmin, Zili Zhang +25 · 15 citations
Engineering · Computer Science · #VLSI and FPGA Design Techniques #Iterative Learning Control Systems #Algorithms and Data Compression
- DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
2024/08/08 by Zili Zhang, Yinmin Zhong, Zhang, Zili +12 · 5 citations
Computer Science · #Distributed #FOS: Computer and information sciences #Natural Language Processing Techniques #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
- Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
2025/06/10 by Ailin Huang, Huang, Ailin, Bingxin Li +143 · 4 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
- ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels
2026/07/20 by Bingyang Wu, Chao Jin, Zili Zhang +6
#cs.DC