Inference economics of language models

We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory bandwidth, network bandwidth and latency constraints; and optimizes over different parallelism setups and batch sizes to find the ones that optimize serial inference speed at a given cost per token. We use the model to compute Pareto frontiers of serial speed versus cost per token for popular language models.

Paper

References (10)

05Massively Scale Your Deep Learning Training with NCCL 2.42019 · NVIDIA Developer Blog
07Quantizing models is most useful when quantization allows the model to fit inside a discontinuous hardware boundary such as a single GPU or a single node
08and as our model would predictgenerally
09NVIDIAgithub
10Our analysis is based strictly on short-context inference in which the attention mechanism is assumed to be negligible, both for arithmetic and memory reads

Similar papers

© 2026 NYSGPT2525 LLC