First of all, thank you for this review and the valuable input. In the following, we address the specific issues you mentioned in the order they were raised and note where we improved the paper accordingly.
**Concerning the potential simplification through geographical locations**
Geographically close stations do not necessarily imply more straightforward causal discovery. First, geographical information is provided but not intended to be used in the causal discovery. Second, The causal effect strength and the elevation differences also play a prominent role. For example, it is very hard to determine the causal effect of a river carrying little water on a large body of water, such as the Elbe River, even if they are geographically close. We revised Section 3.3 (geographical realities) to address your point. Notably, the experiments on graph sets “close 3” and “close 5” (experiment section 1) seem to confirm this intuition as methods typically perform a little (but not much) better on these subgraphs than in comparison to “random 3” and “random 5”.
**Concerning the issue that a focus on a limited subset of a series might be enough for causal discovery**
A specific time series subsection could be used to perform accurate causal discovery. However, selecting the suitable section is non-trivial, as shown in the results of Experiment Set 2. Here, we evaluate some potential candidate subsections but observe no clear-cut improvements. Nonetheless, this is an important point as an appropriate subselection is a way to improve causal discovery algorithms' performance in real-world applications. We are aware of some more advanced strategies that could be used to perform time-series subselection for causal discovery purposes, such as [1] or [2]. However, such strategies have not been a central topic of discussion in the causal discovery literature so far. We added this in the paper's final section as a potential work using our benchmark kit. We also adapted the motivation of Experiment Set 2 to emphasize this point a little more.
**Concerning the focus on a subset of time-series**
We should have noted that more explicitly. The subgraph sets “random 3” and “random 5” include all possible subgraphs with 3 or 5 directly connected nodes, effectively evaluating the complete graph and all available time series. Here, we want to emphasize that this kind of analysis covers the full diversity of the dataset's geographical conditions. Next to this, other sets of subgraphs are used to dissect method performance over different graphical structures. We revised the paper by updating the description (4.1.1) and Table 2.
Further, it would be interesting also to test causal discovery algorithms on larger groups of variables (12+). While this is possible by using the code of our benchmarking kit’s repository, causal discovery algorithms often scale poorly with the number of variables. Thus, we refrain from presenting results on subgraphs with more than 12 nodes.
We hope that we have covered the weaknesses you raised sufficiently. If some uncertainties remain, we are happy to discuss them further.
Finally, we reviewed the text to correct spelling mistakes and improve naming consistency. All the points that you mentioned were corrected. Thank you for pointing them out.
**Concerning your question about the non-time-series settings**
Yes. Other similar benchmarking datasets, especially on non-time-series data, should be feasible. We would be happy to see such work in the future and note this accordingly in the final section.
[1] Ahmad, Wasim, et al. "Deep‐learning based causal inference: A feasibility study based on three years of tectonic‐climate data from Moxa geodynamic observatory." Earth and Space Science 11.10 (2024): e2023EA003430.
[2] Deldari, Shohreh, et al. "Time series change point detection with self-supervised contrastive predictive coding." Proceedings of the Web Conference 2021. 2021.