Response to Post-rebuttal Feedback by Reviewer MYvK
Dear Reviewer MYvK,
Thank you again for your time and effort in reviewing our paper. We appreciate your careful review of our rebuttal materials and your recognition of our efforts in incorporating updates.
We recognize and respect the diverse perspectives regarding the significance of a paper. However, due to the dramatic inconsistency between our opinions, we **kindly request your reconsideration of your reservations on the advancement in performance and technical contributions**. We want to highlight that in accordance with the [NeurIPS 2023 Reviewer Guidelines](https://neurips.cc/Conferences/2023/ReviewerGuidelines), we need specificity, flexibility, and timeliness in your reviews in order for us to better address your concerns.
While we've endeavored with full-time efforts to address your concerns, unfortunately, we've observed that your first feedback logged on 10 hours ago is somewhat vague and limited in specificity and evidence. Here, we will provide further responses to address your concerns, hopefully to your satisfaction, and to help the other reviewers understand the opinions from both sides.
**(1)** Performance advancement
Please kindly refer to our $\underline{\text{Q4 in our rebuttal}}$ above and our further clarification below.
We have made a great effort to support the statistical significance of our improvement and compared our method with typical baselines from previous RL literature, including **DreamerV2 in our main paper, DrQ-v2/Iso-Dream in our supplementary material, and DreamerV3/TransDreamer in our rebuttal**. Our results consistently demonstrate the superior efficacy against these typical RL baselines across various domains and tasks, showing the benefits of our in-the-wild pre-training (IPV) framework and the contribution to the RL community.
Regarding improvement upon our most relevant baseline APV [1] (named as 'IPV w/ vanilla WM' in our paper), it is still statistically significant in Fig. 5b of our paper, aggregated across 48 runs over six tasks of Meta-world. Note that the improvements of 'ContextWM (Pre: O)' against 'ContextWM (Pre: X)' and 'vanilla WM (Pre: O)' are **of a comparable magnitude with improvements made by previous publications** (for example, 'APV (Pre: O / Int: O)' against 'APV (Pre: X / Int: O)' in Fig. 6b of APV paper [1]). While APV makes this improvement with a domain-specific pre-training, we utilize more broadly applicable in-the-wild pre-training.
**(2)** Technical contributions
Please kindly refer to our $\underline{\text{Q2 in our rebuttal}}$ above and we apologize that we did not sufficiently state our contribution in the rebuttal for you.
As stated, **our major technical contribution is to unleash the power of in-the-wild pre-training from videos to boost the sample efficiency of downstream MBRL**. Making world models benefit from in-the-wild pre-training is a critical precondition to scale up to big data and large models since it provides world knowledge widely generalizable and applicable to various downstream tasks. As Reviewer *wPqn* pointed out, 'learning world models on in-the-wild videos is hard', and we highlight that **no previous work has demonstrated positive transfer of a world model from in-the-wild videos** (see Fig. 8c of APV paper [1]). Motivated by the intricate property of in-the-wild contexts, we propose Contextualized World Models, a framework to explicitly separate contextual information and encourage shared dynamics modeling. Our experiments support that **our model successfully breaks the transfer barrier**.
Overall, we have systematically studied **a new problem** (IPV, in-the-wild pre-training from videos), proposed **a new method** tailored for this problem (ContextWM), and demonstrated **significant performance gain** across various domains, which we believe all contribute to the community and help pave the path ahead toward general world models.
[1] Seo, Y., et. al. Reinforcement learning with action-free pre-training from videos. ICML 2022.
We hope that these responses can address your issues and shed light on the significance and solidity of our work. Could you please consider re-evaluating our work based on the updated information? We remain eager to address any lingering concerns and value an open and interactive discussion. Looking forward to your reply.
Best regards,
Authors.