Abstract-In games with stochastic outcomes, evaluating agent performance from limited data is challenging. Results of Monte Carlo sampling do not provide a reliable indicator due to the significant variance. The difficulty of evaluating agents is particularly prominent in mahjong, an incomplete information game with a huge state space. For example, Suphx, which outperformed humans in mahjong, played 5, 760 games against humans in online mahjong to evaluate its performance, which took as long as four months. In this study, we propose MJ-DLVAT, a Deep Learning Value Assessment Technique for Mahjong, which is an evaluation method for mahjong players that provides an unbiased estimate of average ranking with reduced variance. MJ-DLVAT introduces three techniques to manage the extensive game tree and board information in mahjong: splitting the game into subgames, dealing with the variance caused by drawn tiles, dealt tiles and hidden-dora, and introducing neural networks. We created a dataset using online mahjong records and trained a neural network-based value function from scratch. We evaluated MJ-DLVAT on the online mahjong records. We confirmed that the average estimated rankings are unbiased estimators of average ranking and the variance of the estimated ranking is $45.5 \%$ smaller than that of the average ranking. As a result, the number of games required to correctly evaluate a player’s ability is reduced by $45.5 \%$.
Paper
Full text
MJ-DLVAT: A Deep Learning Value Assessment Technique for Mahjong
Semantic Scholar · Computer Science · 2024
Abstract
Abstract-In games with stochastic outcomes, evaluating agent performance from limited data is challenging. Results of Monte Carlo sampling do not provide a reliable indicator due to the significant variance. The difficulty of evaluating agents is particularly prominent in mahjong, an incomplete information game with a huge state space. For example, Suphx, which outperformed humans in mahjong, played 5, 760 games against humans in online mahjong to evaluate its performance, which took as long as four months. In this study, we propose MJ-DLVAT, a Deep Learning Value Assessment Technique for Mahjong, which is an evaluation method for mahjong players that provides an unbiased estimate of average ranking with reduced variance. MJ-DLVAT introduces three techniques to manage the extensive game tree and board information in mahjong: splitting the game into subgames, dealing with the variance caused by drawn tiles, dealt tiles and hidden-dora, and introducing neural networks. We created a dataset using online mahjong records and trained a neural network-based value function from scratch. We evaluated MJ-DLVAT on the online mahjong records. We confirmed that the average estimated rankings are unbiased estimators of average ranking and the variance of the estimated ranking is $45.5 %$ smaller than that of the average ranking. As a result, the number of games required to correctly evaluate a player’s ability is reduced by $45.5 %$.