Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs

Large language models (LLMs) excel in mathematical reasoning and logical problem-solving. Current mainstream paradigms rely on supervised fine-tuning (SFT) or reinforcement learning (RL), yet each possesses distinct strengths: SFT enables the model to internalize fixed patterns, whereas RL empowers it to actively explore and acquire transferable decision-making strategies under reward guidance, thereby enhancing generalization. Recent state-of-the-art methods therefore advocate hybrid schemes, but static switching yields poor cross-task generalization. Inspired by the curriculum learning-quiz mechanism in human reasoning cultivation, we present SASR , a step-wise adaptive hybrid training framework that theoretically unifies SFT and RL and dynamically balances them throughout optimization. SASR begins with an SFT warm-up to establish basic reasoning skills, then employs an adaptive algorithm that monitors gradient norms and divergence from the initial distribution to seamlessly integrate SFT with the online RL method GRPO. By tracking the model’s training status and sequentially adjusting the process, SASR ensures a smooth transition between paradigms, preserving core reasoning while exploring diverse paths. Experimental results demonstrate that SASR outperforms SFT, RL, and static hybrid training methods.

Paper

Similar papers

© 2026 NYSGPT2525 LLC