Exploring Model Compression Techniques for Efficient Inference of Large Language Models

<div> The growing size and complexity of large language models (LLMs) have driven significant advances in natural language processing tasks. However, the extensive computational, memory, and energy demands of these models pose challenges for deployment in resource-constrained environments. Model compression techniques offer a promising solution by reducing model size and inference latency while maintaining high performance. This survey provides a comprehensive overview of the key compression strategies applied to LLMs, including pruning, quantization, knowledge distillation, low-rank factorization, and emerging hybrid approaches. We discuss the theoretical foundations, practical implementations, and trade-offs associated with each method. Additionally, we present prominent evaluation metrics and benchmark datasets used to assess compressed models, ensuring fair and reproducible comparisons. Furthermore, we outline key challenges and promising future research directions, such as adaptive compression techniques, hardware-aware optimizations, and sustainable AI practices. By consolidating recent advancements in model compression, this survey aims to guide researchers and practitioners toward developing efficient LLMs capable of delivering robust performance in diverse real-world applications. </div>

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC