Zheng, Zhen
- DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
2020/07/02 by Fan, Shiqing, Rong, Yi, Meng, Chen +10 · 13 citations
#Distributed #FOS: Computer and information sciences #Parallel #and Cluster Computing (cs.DC)
- FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
2024/01/25 by Haojun Xia, Zhen Zheng, Xia, Haojun +25 · 1 voice · 1 citation
Computer Science · #cs.LG #cs.AI #cs.AR
- ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks
2023/12/14 by Xiaoxia Wu, Haojun Xia, Wu, Xiaoxia +19 · 4 citations
Computer Science · Engineering · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Advanced Neural Network Applications
- BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
2024/11/29 by Zhen Zheng, Zheng, Zhen, Xin Ji +9 · 5 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Data Storage Technologies #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #DNA and Biological Computing #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #and Cluster Computing (cs.DC)
- MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
2024/12/19 by Zheng Zhen, Zheng, Zhen, Chuanjie Liu +2 · 5 citations
Engineering · #Advanced Algorithms and Applications #Advancements in Photolithography Techniques #FOS: Computer and information sciences #Image Processing Techniques and Applications #Machine Learning (cs.LG)
- Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
2023/09/19 by Xia, Haojun, Zheng, Zhen, Li, Yuchao +6 · 2 citations
#Distributed #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Machine Learning (cs.LG) #Parallel #and Cluster Computing (cs.DC)
- FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads
2020/09/23 by Zhen Zheng, Zheng, Zhen, Pengzhan Zhao +15 · 2 citations
Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Parallel Computing and Optimization Techniques #and Cluster Computing (cs.DC)