Enhancing Robustness of Reward Models via Hidden Equilibrium Balancing in Large Language Models
The alignment of large language models (LLMs) with human values through Reinforcement Learning from Human Feedback (RLHF) has become essential in natural language processing. However, existing approaches to optimizing reward models often suffer from overfitting, leading to suboptimal generalization. This paper introduces Hidden Equilibrium Balancing (HEB), a novel regularization technique that maintains a balanced state between text generation capabilities and reward model robustness. HEB operates directly on the hidden states of LLMs and reward models, enforcing both alignment and equilibrium through a principled regularization framework. We present the mathematical formulation of HEB and conduct experiments on the GLUE benchmark and on a preference-based RLHF alignment task using the Anthropic Helpful-Harmless dataset. Our results demonstrate that HEB significantly improves generalization performance, reducing the generalization gap by up to 45% compared to standard regularization methods. The proposed method achieves superior robustness while maintaining computational efficiency, making it suitable for practical RLHF applications.
Paper
Full text
Enhancing Robustness of Reward Models via Hidden Equilibrium Balancing in Large Language Models
Semantic Scholar · 2026
Abstract
The alignment of large language models (LLMs) with human values through Reinforcement Learning from Human Feedback (RLHF) has become essential in natural language processing. However, existing approaches to optimizing reward models often suffer from overfitting, leading to suboptimal generalization. This paper introduces Hidden Equilibrium Balancing (HEB), a novel regularization technique that maintains a balanced state between text generation capabilities and reward model robustness. HEB operates directly on the hidden states of LLMs and reward models, enforcing both alignment and equilibrium through a principled regularization framework. We present the mathematical formulation of HEB and conduct experiments on the GLUE benchmark and on a preference-based RLHF alignment task using the Anthropic Helpful-Harmless dataset. Our results demonstrate that HEB significantly improves generalization performance, reducing the generalization gap by up to 45% compared to standard regularization methods. The proposed method achieves superior robustness while maintaining computational efficiency, making it suitable for practical RLHF applications.