Authors’ response to Reviewer VEfR (Part 1)
Thanks for your time and effort in reviewing our manuscript, as well as for finding the problem highly relevant due to the importance of considering heteroscedastic settings in linear DAG learning. Moreover, we appreciate your valuable suggestions to improve our numerical experiments in several ways, and your request for further elaboration on model identifiability issues. Point-by-point responses to your constructive comments and associated requests for changes follow, which we believe have led to an improved revised paper. We strive to improve our paper and will be happy to continue the scholarly discussion if any lingering issues remain.
**Identifiability.** In the second paragraph of Section 1 we discuss different choices of score functions (likelihood and regression-based) and their relative merits. But with regards to guarantees, you bring up an excellent point that certainly benefits from additional elaboration. In (Loh and Buhlmann, 2014), the overarching assumption is that *the exogenous noise covariance matrix is known* up to a common scale across variables (e.g., in the homoscedastic setting), while CoLiDE also endeavors to estimate the exogenous noise variances. In the homoscedastic setting, the weighted least squares (LS) criterion therein and the CoLiDE-EV score function are equivalent, since the latter is obtained from weighted LS via a constant scaling and shift that will not affect the optimal $\mathbf{W}$ solution. Hence, the identifiability results in (Loh and Buhlmann, 2014) will carry over for both linear Gaussian and non-Gaussian SEMs, again, provided the diagonal noise covariance matrix is known up to a constant factor.
Now, the story is quite different in the more challenging heteroscedastic setting. As we now spell out in the introduction of the revised paper, for general linear Gaussian models the DAG structure is non-identifiable from observational data alone. Interestingly, just like GOLEM (Ng et al, 2020) and for general (non-identifiable) linear Gaussian SEMs, we can show that as $n\to\infty$ CoLiDE-NV probably yields a DAG that is *quasi-equivalent* to the ground truth graph. These guarantees for (heteroscedastic) linear Gaussian SEMs build on the interesting theoretical framework put forth in (Ghassami et al, 2020); see also Appendix C.
Comments along these lines have been included in the revised manuscript; please check the paragraph immediately preceding Section 4.1.
**Data standardization in heteroscedastic settings.** Following your suggestion, we have conducted an additional experiment involving standardized data in a heteroscedastic setting; please check Appendix E.4 in the revised manuscript. Despite an overall degradation in performance of all tested baselines resulting from such standardization (see Figure 9), CoLiDE-NV maintains its status as the superior method when compared to other state-of-the-art approaches.
**Improved metrics over Markov equivalence classes.** This point is well taken. Following your suggestion, we have introduced a new metric termed SHD-C, defined in Appendix D.3 and computed as follows. We initially map both the estimated graph and the ground truth to their respective Completed Partially Directed Acyclic Graphs and subsequently calculate the SHD between them. Notably, SHD-C has been employed in (Ng et al, 2020); which also dealt with (non-identifiable) linear Gaussian SEMs. *SHD-C results are now reported in all relevant tables.* Furthermore, to complement the experiments presented in the main body of the paper, we have included Figures 5 and 8 in Appendix E.
**Gaussian log-likelihood in the heteroscedastic setting.** Thanks again for this constructive feedback. We conducted new experiments specifically aimed at assessing the performance of the CoLiDE score functions in comparison with the Gaussian log-likelihood (as utilized in GOLEM), *by decoupling their effect from the need of a (soft or hard) DAG constraint*. Please check Appendix E.9 in the revised manuscript for a detailed description of the experimental setting. As an ablation study, we believe this is even more informative than augmenting both score functions with, say, the DAGMA acyclicity function. Our results indicate that for smaller DAGs (say with 10 nodes), the log-likelihood score functions outperform concomitant estimators. However, as the number of nodes increases, the score functions utilized in CoLiDE-EV and CoLiDE-NV exhibit superior performance (measured in terms of normalized SHD and SHD-C) when compared to the log-likelihood score functions of GOLEM-EV and GOLEM-NV.