Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
Hybrid training methods for large language models combine supervised fine-tuning (SFT) on expert demonstrations with reinforcement learning (RL) on model rollouts, typically at sample granularity. We propose Entropy-Gated Selective Policy Optimization (EG-SPO), a three-stage framework that extends sample-level mixing with token-level gradient modulation. Stage 1 (SFT Expert Learning) establishes a reliable warm-up policy using expert demonstrations with pure SFT loss. Stage 2 (RL Rollout Generation) generates model rollouts from the current policy and computes per-token predictive entropy. Stage 3 (EG-SPO Main Mechanism) applies entropy-gated gradient allocation: a Predictive Entropy Module routes high-entropy tokens to full PPO updates (encouraging exploration) and low-entropy tokens to ϕ-attenuated PPO (reducing variance, preserving knowledge). Critically, both branches incorporate the advantage function At, ensuring that incorrect trajectories receive consistent negative learning signals—preventing reinforcement of confident errors. Our method achieves consistent improvements on mathematical reasoning benchmarks: +3.8% on AIME and +2.9% on MATH over the CHORD-ϕ baseline, with only 3.4% computational overhead.
Paper
References (17)
Scroll for more · 5 remaining