Integrating Knowledge Distillation into AlphaGo Zero for Gomoku Board Game

AlphaGo Zero achieves strong performance in complex board games through tabula rasa self-play, but its lengthy training process poses challenges for resource-constrained environments, limiting its accessibility. This study aims to address this issue by integrating knowledge distillation (KD) into the AlphaGo Zero framework to optimize Gomoku model training. An offline KD strategy was adopted: a <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{2 0 0}$</tex>-hour pre-trained model served as the teacher, providing soft targets, while the student model learned from both these soft targets and its own self-play hard targets via a hybrid loss function with time-decaying weights. Experiments show the KD model converges 1500 steps earlier than non-KD models. The student matches the teacher at 900 training steps, surpasses it at 1000 steps, and eventually integrates teacher knowledge with novel strategies. Key enablers are informative soft target guidance and dynamic weights balancing teacher reliance and exploration. Limitations include teacher dependency and static hyperparameters. This framework offers an efficient reinforcement learning solution for board games.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC