Rong Gu
- SSV: Sparse Speculative Verification for Efficient LLM Inference
2026/05/19 by Zhibin Wang, Ziyu Zhong, Nuo Shen +3 · 2 voices
#cs.OS
- Echo: Efficient Co-Scheduling of Hybrid Online-Offline Tasks for Large Language Model Serving
2025/03/01 by Zhibin Wang, Shipeng Li, Wang, Zhibin +16 · 2 citations
Computer Science · #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Big Data and Digital Economy #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #and Cluster Computing (cs.DC)
- Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding
2025/08/29 by Zhibin Wang, Zhonghui Zhang, Wang, Zhibin +18 · 1 citation
Computer Science · #Target Tracking and Data Fusion in Sensor Networks #Anomaly Detection Techniques and Applications #Distributed Sensor Networks and Detection Algorithms
- Revisiting Service Level Objectives and System Level Metrics in Large Language Model Serving
2024/10/18 by Zhibin Wang, Shipeng Li, Wang, Zhibin +14 · 1 citation
Computer Science · Engineering · #Artificial Intelligence (cs.AI) #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Machine Learning (cs.LG) #Power Systems and Technologies #Service-Oriented Architecture and Web Services
- SpecLA: Efficient Speculative Decoding for Linear-Attention Models
2026/07/18 by Zhibin Wang, Xuying Han, Zhaohua Yang +5
#cs.CL