With regards to our assessment of AlphaZero, we could have been clearer in the previous answer. It doesn't saturates after only 100 iterations, but around 400. 'with marginal improvements after 100 steps' was a poorly worded way of saying that the model stayed significantly worse than AlphaGateau, eventually reaching a rating between 300 and 500 Elo, compared to the more than 1500 of AlphaGateau. Seeing these results, we decided to only train AlphaGateau for 100 iterations as it seemed that the behavior of the model during these early iterations was the most important part, with both AlphaZero and AlphaGateau increasing significantly less afterwards.
However, we do now agree with your criticism that the asymptotic behavior is also important to include. We will include in the revised version of the paper 500 steps of training of AlphaZero. We are currently running this experiment with our up-to-date code, with Elo ratings going up to iteration 300, where it still only goes from 100 elo at iteration 100 to around 450 by iteration 300.
We will also note that AlphaGateau models don't saturate either by iteration 100. This was not clear in our initial figures as we only evaluated Elo rating once every 5 iterations, but it is clearer in our new figures evaluated every second iteration, such as the one in the global rebuttal pdf. We would further like to extend the runs of the AlphaGateau models to at least iteration 200, and will include those new runs in the revised paper.
Yes, we only updated the data generating model once every step, and chose how many games to generate in one step to saturate our GPU. We could generate less games per step, but we would waste computing resources by doing so. We experimented trying to run more than one full epoch between each step of generation, but didn't notice any significant improvement behind the first 10 steps, so we decided to stick to only one epoch per step (and around 500-1000 batches) as the training already took the majority of our compute time.
For the details of the training of AlphaZero, we generated 256 games at each step, and used batches of 2048 positions, for a total of 488 batches each step (except the first 6-7 steps, while the frame buffer is filling up). Each set of 256 generated games used newly updated network parameters. We also only used 128 MCTS simulations in each of our experiments.
We didn't attempt to anchor our Elo ratings to other real agents, as our main results are the improvements of our architecture when compared to its AlphaZero basis. We however did let our latest fine-tuned for 20 steps 6-layer AlphaGateau model play around 600 blitz and bullet games on lichess, against mostly other bots, ending with an approximate Elo of 1800 in blitz, and 2000 in bullet. However, we had to adjust the number of MCTS simulations in order to make our model take an appropriate mostly constant amount of time to evaluate each move, which was often different to the number used during training, so these ratings are only indicative of the approximate abilities of that model.