FastFormers: Highly Efficient Transformer Models for Natural Language Understanding

Transformer-based models are the state-of-the-art for Natural Language\nUnderstanding (NLU) applications. Models are getting bigger and better on\nvarious tasks. However, Transformer models remain computationally challenging\nsince they are not efficient at inference-time compared to traditional\napproaches. In this paper, we present FastFormers, a set of recipes to achieve\nefficient inference-time performance for Transformer-based models on various\nNLU tasks. We show how carefully utilizing knowledge distillation, structured\npruning and numerical optimization can lead to drastic improvements on\ninference efficiency. We provide effective recipes that can guide practitioners\nto choose the best settings for various NLU tasks and pretrained models.\nApplying the proposed recipes to the SuperGLUE benchmark, we achieve from 9.8x\nup to 233.9x speed-up compared to out-of-the-box models on CPU. On GPU, we also\nachieve up to 12.4x speed-up with the presented methods. We show that\nFastFormers can drastically reduce cost of serving 100 million requests from\n4,223 USD to just 18 USD on an Azure F16s_v2 instance. This translates to a\nsustainable runtime by reducing energy consumption 6.9x - 125.8x according to\nthe metrics used in the SustaiNLP 2020 shared task.\n

Paper

References (30)

Scroll for more · 18 remaining

Similar papers

© 2026 NYSGPT2525 LLC