Summary
This work presents a new semantic map representation combined with leveraging a planning policy which uses multiple prior methods in a single system to explore and reach semantic targets in novel environments. The correctness of the 2D semantic map is measured as well as semantic navigation performance in novel environments in simulation.
Strengths
The authors do a good job of performing ablations of their own method to determine the relative contribution of each of the components of their pipeline. The diagrams of their proposed pipeline are well made and clear. The description of the approach is also presented clearly and the motivation and details explained well.
Weaknesses
The experimental results have substantial errors. The proposed method, SkillTron+SegmATRon, is evaluated on a manually selected subset of validation episodes from the HM3D dataset. Then, numbers from the Habitat Challenge are input in comparison which evaluate on a completely different set of test episodes which are not public. SkillTron+SegmATRon is not accurately compared (on one shared test set!) against published state-of-the-art methods for semantic navigation.
If the authors want to have a 1-1 comparison with alternate approaches without manually rerunning baselines on their custom testset, they should evaluate SkillTron+SegmATRon on the full test set for HM3D (the dataset they use) and then numbers from other papers which evaluate on this standard and public test set can be directly input into the comparison table. If they would like to evaluate against methods which run in the Habitat Challenge, they should submit their code to the Habitat Challenge so that they have also run on the same secret test set. However, running on the challenge should not replace running on a standard publicly available test (i.e. the HM3D test set) so that their paper has publicly reproducible evaluation results.
Also, the authors claim “significant improvement” in the experiment section of SkillTron+SegmATRon over SkillTron+OneFormer (which seem to have been evaluated on the same test set) when the change in both success rate, SPL, and SoftSPL is 0.02 or less in every case and no standard error metrics are reported. So, it is not clear in the ablations whether the semantic map method actually yields a statistically significant change in performance at all.
In addition to the experimental errors, there are many statements throughout the paper in comparisons to related work which also are conjectures stated as facts. For example, in the introduction “such approaches are not correctly linked to the algorithms for planning map trajectories” regarding the approach in (Ramakrishnan et al., 2022). Why is the RL approach used to plan a map trajectory in this paper “incorrect”?
Also, the authors should consider in the presentation of their method motivation: If the goal of their work is to find a highly accurate semantic map of the observed area, why should highly performant semantic SLAM methods not be used?
Questions
Note: the correct review template is not used. The heading says “Published as a conference paper at ICLR 2024.” instead of “Under review as a conference paper at ICLR 2024”. The authors are using the camera ready template with the author names deleted - not the anonymized submission template.
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.