Wang, Zeke
- LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
2024/03/11 by Changyue Liao, Mo Sun, Liao, Changyue +13 · 1 voice · 8 citations
Computer Science · Medicine · #Distributed #Distributed and Parallel Computing Systems #FOS: Computer and information sciences #Medical Imaging Techniques and Applications #Parallel #Parallel Computing and Optimization Techniques #and Cluster Computing (cs.DC) #cs.DC
- Demystifying Datapath Accelerator Enhanced Off-path SmartNIC
2024/02/05 by Xuzheng Chen, Jie Zhang, Chen, Xuzheng +18 · 4 citations
Computer Science · Engineering · #68M10 #C.2.1 #Embedded Systems Design Techniques #FOS: Computer and information sciences #Networking and Internet Architecture (cs.NI) #Smart Grid Security and Resilience
- DeFT: Decoding with Flash Tree-attention for Efficient Tree-structured LLM Inference
2024/03/30 by Yao, Jinwei, Chen, Kaiqi, Zhang, Kexun +4 · 3 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
- Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN Training
2023/05/26 by Jie Sun, Li Su, Sun, Jie +19 · 2 citations
Computer Science · #Advanced Graph Neural Networks #Caching and Content Delivery #Distributed #FOS: Computer and information sciences #Machine Learning and ELM #Parallel #and Cluster Computing (cs.DC)
- Benchmarking High Bandwidth Memory on FPGAs
2020/05/09 by Wang, Zeke, Huang, Hongjing, Zhang, Jie +1 · 1 citation
#FOS: Computer and information sciences #Hardware Architecture (cs.AR)
- MARS: Exploiting Multi-Level Parallelism for DNN Workloads on Adaptive Multi-Accelerator Systems
2023/07/23 by Guan Shen, Shen, Guan, Jieru Zhao +13 · 1 citation
Computer Science · Engineering · Neuroscience · #Advanced Memory and Neural Computing #Advanced Neural Network Applications #Artificial Intelligence (cs.AI) #Brain Tumor Detection and Classification #Distributed #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Parallel #and Cluster Computing (cs.DC)
- Helios: An Efficient Out-of-core GNN Training System on Terabyte-scale Graphs with In-memory Performance
2023/10/02 by Jie Sun, Mo Sun, Sun, Jie +15 · 1 citation
Computer Science · Engineering · #Advanced Graph Neural Networks #Graph Theory and Algorithms #Ferroelectric and Negative Capacitance Devices
- TorchGT: A Holistic System for Large-scale Graph Transformer Training
2024/07/19 by Zhang, Meng, Sun, Jie, Hu, Qinghao +4 · 1 citation
#Artificial Intelligence (cs.AI) #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #and Cluster Computing (cs.DC)