vix.ing · top · new · best · stats · spec

Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems

2024/07/09 by Amey Agrawal, Agrawal, Amey, Anmol Agarwal +13 · 2 citations
Computer Science · Decision Sciences · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Data Mining Algorithms and Applications #Data Quality and Management #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Parallel #Semantic Web and Ontologies #and Cluster Computing (cs.DC)

paper · pdf · doi:10.48550/arxiv.2407.07000

openalex publication_date 2024/07/09 · openalex created_date 2024/07/11 · openalex updated_date 2026/07/28

Abstract

Serving large language models (LLMs) in production can incur substantial costs, which has prompted recent advances in inference system optimizations. Today, these systems are evaluated against conventional latency and throughput metrics (eg. TTFT, TBT, Normalised Latency and TPOT). However, these metrics fail to fully capture the nuances of LLM inference, leading to an incomplete assessment of user-facing performance crucial for real-time applications such as chat and translation. In this paper, we first identify the pitfalls of current performance metrics in evaluating LLM inference systems. We then propose Etalon, a comprehensive performance evaluation framework that includes fluidity-index -- a novel metric designed to reflect the intricacies of the LLM inference process and its impact on real-time user experience. Finally, we evaluate various existing open-source platforms and model-as-a-service offerings using Etalon, discussing their strengths and weaknesses. Etalon is available at https://github.com/project-etalon/etalon.

Cited by

Related