Dear authors, thank you for the follow-up response. I will use this comment to summarize my position now that I believe I have a better understanding of the manuscript.
**Motivation**
First, I appreciate you rephrasing the hyperbolic language. It makes the claims sound much more precise and fosters clarity of what is provided.
That being said, I think that the scalability and interactivity arguments are not fleshed our sufficiently. For instance, [1 mine] also have an interactive world model that is scalable which is probably worth citing. What differentiates the presented work here is that the world-model operates in pixel-observation space. This in itself is interesting because I am not aware of any work that has done this on this complexity. The second argument that the rebuttal makes is that **one** model for downstream fine-tuning is what is being provided. However, this point does currently not play a central role in the writing.
Overall, I think the paper would benefit strongly from re-focusing the main story on pixel-observation spaces and a single-model for downstream learning rather than the interactive aspect. I believe the interactive and scalable aspects are not necessarily what makes this work stand out (I personally would adjust the title to include pixel-spaces/observation spaces. Note that that does not mean the name of the model needs to change since interactive video model is still true.). A single, pixel observation-based world model for easy downstream finetuning would fit the story of this work better.
**Tokenizer**
I appreciate all the clarifications on this. Now that I understand better, I believe the text again would benefit from re-writing, focusing on the need of this tokenizer to enable embedding actions seemlessly while also providing an efficient encoding similar to what is needed for video prediction in iVideoGPT rather than things like context-lengths. This would tie the introduction of the tokenizer much better to the objective of building a world model which is currently missing.
**Experiments**
Thank you for running the additional experiments. I think the experiment section of this paper is going to be quite extensive which is great. Overall, I believe my point about the strengths of the experiments remains. Yet, given my previous explanation we have to put them in a different light. A world model in pixel observation space that is equally good as latent space models is a decent contribution. Tying the language of the claims to the experiments is what remains.
One crucial addition are the additional dreamer experiments to strengthen the claim that previous models are not scalable. However, as I pointed out, Dreamer is not the model that claims to be scalable, but rather [1] does. I understand that there is no time in the rebuttal to consider a comparison here but I do think that comparison would make this paper a stronger submission. Since there is no time for this now, I will not include this into my recommendation to the author's disadvantage but rather provide this as feedback for the next iteration of the paper.
"However, as shown in Fig. 6 of the paper, data scaling does provide benefits---pretraining iVideoGPT enhances MBRL performance."
Figure 6 only demonstrates that pretraining is useful which is something we have known from offline RL, not that we need a lot of fine-tuning data. There are no experiments with varying amounts of data for **down-stream** performance. In Figure 8b the paper demonstrates that validation loss of pre-training goes down with larger models but in the rebuttal it's stated that that has no impact on the down-stream performance on meta-world. Thus, I believe there is still a disconnect between analyzing the scalability (both model and data) and down-stream performance of the model.
Thank you also for clarifying training times, I underestimated the reduction in computational cost from your tokenizer.
**Summary**
Overall, I think this paper makes a decent contribution and others might benefit from it being published. However, I think there are various sections that I personally would recommend re-writing. If this was a journal paper, it would be an accept with revisions. If we had more time, I would recommend borderline acceptance and raise my score to 5. Unfortunately, we don't have this option here which puts us at an impasse. Thus I am going to err on the side of optimism and trust that the authors incorporate large parts of my feedback and I will recommend acceptance. Most of the issues that I have with the paper can be fixed by (in some parts minor) re-writing. I am changing my scores as follows:
Presentation: 3 -> 2
Contribution: 2 -> 4
Overall score: 3 -> 6
[1] TD-MPC2: Scalable, Robust World Models for Continuous Control. Nicklas Hansen, Hao Su, Xiaolong Wang. ICLR 2024.