Panwar, Ashish
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
2024/03/04 by Amey Agrawal, Agrawal, Amey, Nitin Kedia +13 · 151 citations
Computer Science · #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #and Cluster Computing (cs.DC)
- PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch
2025/03/25 by Abhishek Ghosh, Ajay Nayak, Ghosh, Abhishek +5 · 9 voices · 2 citations
Computer Science · #cs.LG
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
2023/08/31 by Amey Agrawal, Ashish Panwar, Agrawal, Amey +9 · 53 citations
Computer Science · Engineering · #Distributed #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
- Vidur: A Large-Scale Simulation Framework For LLM Inference
2024/05/08 by Amey Agrawal, Agrawal, Amey, Nitin Kedia +13 · 27 citations
Decision Sciences · #Simulation Techniques and Applications
- vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
2024/05/07 by Ramya Prabhu, Prabhu, Ramya, A.K. Nayak +7 · 25 citations
Computer Science · #Advanced Data Storage Technologies #Cryptography and Data Security #FOS: Computer and information sciences #Machine Learning (cs.LG) #Operating Systems (cs.OS) #Security and Verification in Computing
- Mitosis: Transparently Self-Replicating Page-Tables for Large-Memory Machines
2019/10/11 by Achermann, Reto, Panwar, Ashish, Bhattacharjee, Abhishek +2 · 1 citation
#FOS: Computer and information sciences #Hardware Architecture (cs.AR) #Operating Systems (cs.OS) #Performance (cs.PF)