Summary
This paper developed a state representation learning method leveraging an unbalanced atlas (UA). The authors have modified the ST-DIM algorithm to align with the proposed UA paradigm. Although the main contribution is not stated intuitively, empirical evaluations on 19 games of the AtariARI benchmark suggested an improved performance compared with three established baseline methods (many existing self-supervised methods are omitted for comparison). Furthermore, the authors performed a comprehensive ablation study for the design choices of the proposed method.
Strengths
+ The experiments are conducted across 19 games of the AtariARI benchmark, covering a variety of vision tasks.
+ There are comprehensive ablation studies for the technical components of the proposed method.
Weaknesses
- The clarity of the introduction could be enhanced by providing a more explicit context for the specialized terminology introduced (see Q1-3).
- The comparison would benefit from the inclusion of key baseline models which are currently absent (see Q4).
- Tables 1 and 2 appear to be redundant, presenting analogous results through different evaluative metrics (F1 score and Accuracy, respectively). Although a comprehensive evaluation is encouraged, putting these two sizable tables back to back in the main paper gives the impression of lacking sufficient materials for the paper. It would be more appropriate to consolidate these findings, perhaps through a combined analysis or in supplementary materials, to avoid repetition and maintain the conciseness of the paper.
- the paper lacks a clear statement of its underlying motivation and significance, which is pivotal for readers to comprehend the value and potential impact of the research (see Q5).
Questions
1. The introduction used specialized terminology that may not be universally familiar, necessitating additional clarification for a broader audience. Specifically, the first sentence of the third paragraph introduces concepts such as *manifold*, *atlas*, *local structure*, and *chart*, which would benefit from further exposition to contextualize the study and its objectives.
2. The paper's motivation remains unclear, partly owing to the use of undefined terms. The concept of an *atlas*, and particularly the distinction between *unbalanced* and *balanced* atlases within this framework, needs clarification. The terms *prior distribution* and *membership probability distribution* introduced later also lack clear definitions, impeding the reader's understanding.
3. In the 3rd paragraph of Section 4, *d* and *n* are used without proper definition.
4. This paper suggests that pre-training a model using a reinforcement learning task and then fine-tuning it on downstream reinforcement learning tasks is beneficial. However, this point is not fully demonstrated because the authors did not compare the proposed method with other self-supervised learning methods, as reviewed in the introduction, e.g., contrastive models (SimCLR is compared) and generative models (none is compared).
5. From the introduction section, it is not intuitive to me why this study is important. For example, the last paragraph lists technical achievements but does not convey their broader significance. Specifically, (1) fitting the ST-DIM to UA paradigm (why UA paradigm is important?), (2) detailed ablations for better design choices (that's standard, not sure if it counts as a contribution), and (3) representing a manifold with a larger number (why this is important anyway?). Not limited to the introduction section, the authors did not describe the significance of the proposed method and all these ablation studies in the entire paper.
Rating
6: marginally above the acceptance threshold
Confidence
2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.