Ashish Panwar
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
2024/03/04 by Amey Agrawal, Nitin Kedia, Agrawal, Amey +13 · 131 citations
Computer Science · #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #and Cluster Computing (cs.DC)
- PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch
2025/03/25 by Abhishek Ghosh, Ajay Nayak, Ghosh, Abhishek +5 · 9 voices · 2 citations
#cs.LG
- SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
2023/08/31 by Amey Agrawal, Agrawal, Amey, Ashish Panwar +9 · 45 citations
Computer Science · Engineering · #Distributed #FOS: Computer and information sciences #Ferroelectric and Negative Capacitance Devices #Machine Learning (cs.LG) #Natural Language Processing Techniques #Parallel #Topic Modeling #and Cluster Computing (cs.DC)
- Vidur: A Large-Scale Simulation Framework For LLM Inference
2024/05/08 by Amey Agrawal, Agrawal, Amey, Nitin Kedia +13 · 24 citations
Decision Sciences · #Simulation Techniques and Applications
- vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
2024/05/07 by Ramya Prabhu, A.K. Nayak, Prabhu, Ramya +7 · 23 citations
Computer Science · #Advanced Data Storage Technologies #Cryptography and Data Security #FOS: Computer and information sciences #Machine Learning (cs.LG) #Operating Systems (cs.OS) #Security and Verification in Computing