Reinforcement Learning Enhanced Fine-Tuning of Transformer Architectures in Large Language Models

Transformer-based LLMs like GPT and BERT have excelled in many NLP tasks, traditional full-parameter fine-tuning is computationally costly and inefficient when it comes to tailoring these models to domain-or task-specific goals. Current methods try to overcome these limitations, but they have issues with scalability, training instability, dependence on expensive human feedback, limited expressiveness, and weak long-horizon optimization. Some examples of these approaches are Low-Rank Adaptation (LoRA), Decision Transformers, and Reinforcement Learning from Human Feedback (RLHF). A reinforcement learning enhanced fine-tuning architecture integrating parameter-efficient adaptation modules is proposed in this research to address these challenges. The suggested method permits scalable, efficient, and stable model adaptation by freezing pretrained transformer parameters and optimizing lightweight fine-tuning components with task-level reward signals. The suggested method outperforms supervised fine-tuning in terms of recall, task accuracy, F1-score, and reward optimization; moreover, it maintains stable convergence and reduces computational overhead, making it an excellent choice for large-scale and practical applications.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC