vix.ing · top · new · best · stats · spec

Phanishayee, Amar

  1. Efficient Large-Scale Language Model Training on GPU Clusters Using\n Megatron-LM
    2021/04/09 by Deepak Narayanan, Mohammad Shoeybi, Narayanan, Deepak +21 · 120 citations
    Computer Science · #Computation and Language (cs.CL) #Distributed #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
  2. PipeDream: Fast and Efficient Pipeline Parallel DNN Training
    2018/06/08 by Aaron Harlap, Deepak Narayanan, Harlap, Aaron +11 · 34 citations
    Computer Science · #Advanced Neural Network Applications #Distributed #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Parallel #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
  3. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads
    2019/01/17 by Myeongjae Jeon, Jeon, Myeongjae, Shivaram Venkataraman +9 · 28 citations
    Computer Science · #Cloud Computing and Resource Management #Distributed #FOS: Computer and information sciences #IoT and Edge/Fog Computing #Parallel #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
  4. Memory-Efficient Pipeline-Parallel DNN Training
    2020/06/16 by Narayanan, Deepak, Phanishayee, Amar, Shi, Kaiyu +2 · 19 citations
    #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Parallel #and Cluster Computing (cs.DC)
  5. Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads
    2020/08/20 by Deepak Narayanan, Narayanan, Deepak, Keshav Santhanam +7 · 18 citations
    Computer Science · #Cloud Computing and Resource Management #Distributed #FOS: Computer and information sciences #Parallel #Parallel Computing and Optimization Techniques #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
  6. Themis: Fair and Efficient GPU Cluster Scheduling
    2019/07/02 by Kshiteej Mahajan, Mahajan, Kshiteej, Arjun Balasubramanian +11 · 12 citations
    Computer Science · #Cloud Computing and Resource Management #Distributed #FOS: Computer and information sciences #IoT and Edge/Fog Computing #Parallel #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
  7. DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
    2024/03/04 by Strati, Foteini, Mcallister, Sara, Phanishayee, Amar +2 · 13 citations
    #Distributed #FOS: Computer and information sciences #Parallel #and Cluster Computing (cs.DC)
  8. Blink: Fast and Generic Collectives for Distributed ML
    2019/10/11 by Guanhua Wang, Shivaram Venkataraman, Wang, Guanhua +9 · 8 citations
    Computer Science · #Advanced Neural Network Applications #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Parallel Computing and Optimization Techniques #Stochastic Gradient Optimization Techniques #and Cluster Computing (cs.DC)
  9. Analyzing and Mitigating Data Stalls in DNN Training
    2020/07/14 by Mohan, Jayashree, Phanishayee, Amar, Raniwala, Ashish +1 · 5 citations
    #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Operating Systems (cs.OS) #Parallel #and Cluster Computing (cs.DC)
  10. Efficient Algorithms for Device Placement of DNN Graph Operators
    2020/06/29 by Jakub Tarnawski, Amar Phanishayee, Tarnawski, Jakub +7 · 4 citations
    Computer Science · Engineering · #Advanced Memory and Neural Computing #Advanced Neural Network Applications #Distributed #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Parallel #and Cluster Computing (cs.DC)
  11. Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
    2020/06/05 by Hongyu Zhu, Zhu, Hongyu, Amar Phanishayee +3 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Parallel #Parallel Computing and Optimization Techniques #Performance (cs.PF) #Software Testing and Debugging Techniques #and Cluster Computing (cs.DC)
  12. A Study on the Intersection of GPU Utilization and CNN Inference
    2022/12/15 by Kosaian, Jack, Phanishayee, Amar · 1 citation
    #FOS: Computer and information sciences #Machine Learning (cs.LG) #Performance (cs.PF)
  13. Blox: A Modular Toolkit for Deep Learning Schedulers
    2023/12/19 by Agarwal, Saurabh, Phanishayee, Amar, Venkataraman, Shivaram · 2 citations
    #Distributed #FOS: Computer and information sciences #Parallel #and Cluster Computing (cs.DC)
  14. Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
    2024/04/23 by Adnan, Muhammad, Phanishayee, Amar, Kulkarni, Janardhan +2 · 1 citation
    #Distributed #FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Parallel #and Cluster Computing (cs.DC)
  15. Integrated Hardware Architecture and Device Placement Search
    2024/07/18 by Irène Wang, Wang, Irene, Jakub Tarnawski +5 · 1 citation
    Computer Science · Engineering · #VLSI and Analog Circuit Testing #Embedded Systems Design Techniques #VLSI and FPGA Design Techniques