Summary
The paper proposes a data with a combination of synthetic data and real data of windows, to facilitate research in methods including segmentation and domain adaptation. The paper includes exhaustive description on the specifics of the proposed dataset, as well as benchmark on segmentation under various training/testing strategies/splits to demonstrate the domain gap between real/synthetic splits and splits of different geo locations. The paper also evaluates the impact of synthetic dataset variations on segmentation, including rendering/design/labeling choices.
Strengths
[1] The paper proposes a dataset with both synthetic/real splits on a specific domain, i.e. windows, with high quality renderings and real captures, as well as detailed statistics, benchmarks on the task of segmentation.
[2] The paper elaborately justifies the choice of windows as the domain to model in the proposed dataset, based on its intra-class variation across geographical locations and appearances, availability of source images, and suitability for tasks including segmentation and domain adaptation.
Weaknesses
[1] Limited contribution. The proposal of a new dataset including both synthetic and real data is always appreciated; however, the paper does not fully justify its significance over other existing datasets. For instance, with respect to the application of segmentation, there are a wide range of datasets especially on city scenes or driving scenes that have been extensively utilized for the task and even in the context of domain adaptation. In this paper the choice of windows as the domain is properly justified, but it did not answer the question on whether existing datasets can already serve the same purpose (or more). In this case, the significance of the dataset is in question.
[2] Limited evaluation. The paper evaluates the proposed dataset in segmentation tasks on various settings. However a missed opportunity given the availability of both real and simulated splits and the discovery of domain gap in its evaluation, is to test the dataset on domain adaptation methods. The paper could have tested the dataset with existing domain adaptation methods on what the baseline benchmark looks like, potential advantage or challenges of the dataset when compared to other existing datasets in the context of domain adaptation. Otherwise it is difficult to evaluate the impact of the proposed dataset on domain adaptation research, without including actual evaluation results.
[3] Writing. The paper includes lengthy introduction and justification of the choice of windows, in the Introduction section. The text mostly served its purpose. However the writing can be greatly improved. For instance, reorganizing the first paragraph of the Introduction section from 6 points into an organic flow of arguments can be made: the main task the dataset is motivated towards, issues with existing datasets, features of the proposed dataset which significantly differentiates itself from existing ones, and other major features of the proposed dataset. Additionally, the 6 points can overlap: 1st and 4th points both talk about why choosing the domain of windows, and also overlaps with the entirety of the second paragraph. The 6th point is confusing: any dataset comprised of images can be used to extract image features. Finally, additional introduction of Wood et al. before referring to it may benefit the readers in better understanding the last paragraph of the Related Works section and Label Adaptation of Section 6.
Questions
[1] Can existing domain adaptation methods (e.g. GeoNet and related methods: https://tarun005.github.io/GeoNet/) be applied to the proposed dataset, and what are the observations compared to being applied to other existing datasets?
[3] Can other applications be enabled by the proposed dataset besides segmentation?
Rating
3: reject, not good enough
Confidence
3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked.
Ethics concerns
A team of photographers were hired to take photos of windows in the streets; as a result there is a good chance that some of the original photos may be related to privacy concerns especially those taken of private residences. The paper mentions that such samples are properly filtered in the final dataset; however more details are to be inspected: are the privacy-concerning raw images properly destroyed? Do any such photos end up in the final dataset and how to guarantee that there are none?