Large Language Models (LLMs) have transformed natural language processing by exhibiting advanced capabilities in comprehension, reasoning, and context-aware text generation. However, the rapid escalation in their size and computational requirements presents major obstacles in training, deployment, and energy sustainability. The substantial hardware, memory, and power demands of these models restrict their accessibility and limit real-time or edge-based implementation. To address these concerns, current research emphasizes model-compression and optimization techniques that maintain accuracy while improving efficiency. Methods such as quantization, pruning, knowledge distillation, parameter sharing, and low-rank adaptation have demonstrated promising results in reducing redundancy and accelerating inference. Additionally, hybrid precision and adaptive computation frameworks seek to balance performance with computational cost. This survey consolidates recent progress in efficiency-oriented LLM research, examining their comparative advantages, limitations, and practical trade-offs.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex