vix.ing · top · new · best · stats · spec

Zheng, Zhen

  1. DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
    2020/07/02 by Fan, Shiqing, Rong, Yi, Meng, Chen +10 · 13 citations
    #Distributed #FOS: Computer and information sciences #Parallel #and Cluster Computing (cs.DC)
  2. FP6-LLM: Efficiently Serving Large Language Models Through FP6-Centric Algorithm-System Co-Design
    2024/01/25 by Haojun Xia, Zhen Zheng, Xia, Haojun +25 · 1 voice · 1 citation
    Computer Science · #cs.LG #cs.AI #cs.AR
  3. ZeroQuant(4+2): Redefining LLMs Quantization with a New FP6-Centric Strategy for Diverse Generative Tasks
    2023/12/14 by Xiaoxia Wu, Haojun Xia, Wu, Xiaoxia +19 · 4 citations
    Computer Science · Engineering · #Topic Modeling #Ferroelectric and Negative Capacitance Devices #Advanced Neural Network Applications
  4. BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
    2024/11/29 by Zhen Zheng, Zheng, Zhen, Xin Ji +9 · 5 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Data Storage Technologies #Algorithms and Data Compression #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #DNA and Biological Computing #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #and Cluster Computing (cs.DC)
  5. MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
    2024/12/19 by Zheng Zhen, Zheng, Zhen, Chuanjie Liu +2 · 5 citations
    Engineering · #Advanced Algorithms and Applications #Advancements in Photolithography Techniques #FOS: Computer and information sciences #Image Processing Techniques and Applications #Machine Learning (cs.LG)
  6. Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
    2023/09/19 by Xia, Haojun, Zheng, Zhen, Li, Yuchao +6 · 2 citations
    #Distributed #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Machine Learning (cs.LG) #Parallel #and Cluster Computing (cs.DC)
  7. FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads
    2020/09/23 by Zhen Zheng, Zheng, Zhen, Pengzhan Zhao +15 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Adversarial Robustness in Machine Learning #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Parallel Computing and Optimization Techniques #and Cluster Computing (cs.DC)