HeadTile: A Scalable and Efficient Accelerator for Large Language Model Inference with 3D Memory Integration

Large Language Models (LLMs) inference have gained popularity over the past two years, driven by their high performance achieved through rapid increases in the number of parameters. The inference process of LLMs consists of two distinct stages: prefill and decode, each with unique computational characteristics. While existing neural network inference platforms, such as Google's TPU, perform well during the prefill stage, they often suffer from poor resource utilization during the decode stage. To address this challenge, we propose the scalable Headtile architecture, specifically designed to improve hardware resource utilization. By analyzing the inference behavior of LLMs, we examine how each layer executes on TPUv3 and introduce the Maarg paradigm in Headtile for inter-layer scheduling and mapping. Experimental results show that Headtile can achieve up to 24 × higher throughput in the decode stage compared to TPUv3. In addition, the Maarg paradigm reduces memory accesses by up to 60 % during the prefill stage.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC