vix.ing · top · new · best · stats · spec

Inference economics of language models

2025/06/05 by Ege Erdil, Erdil, Ege · 5 citations
Computer Science · #Distributed #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning and Algorithms #Natural Language Processing Techniques #Parallel #Parallel Computing and Optimization Techniques #and Cluster Computing (cs.DC)

paper · pdf · doi:10.48550/arxiv.2506.04645

openalex publication_date 2025/06/05 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory bandwidth, network bandwidth and latency constraints; and optimizes over different parallelism setups and batch sizes to find the ones that optimize serial inference speed at a given cost per token. We use the model to compute Pareto frontiers of serial speed versus cost per token for popular language models.

Cited by

Related