Dear reviewer, thank you for your comment. We are glad to see that our response made you feel more positive about our work.
* Regarding the elbow method, the idea can be traced back to [1]. We are including a few other references that use or explain this heuristic [2,3,4]. We should point out that this elbow approach was merely tested as an alternative to the threshold approach, and its inclusion in our work was based on the reasonable performance we observed using this method.
* As per your comment regarding "some rigorous results", we would like to remind you that our theory on the identifiability of mechanism shifts is rigorous and all the proofs are given in the appendix. Since our theory is given at the population level, we rely on algorithmic approaches or having a hyperparameter (the threshold $t$) in order to implement our method in finite samples. This is not unique to our work, in fact, even Rolland et al.'s SCORE method uses a heuristic to choose a leaf from a set of nodes---namely, pick the node as the diagonal entry with the lowest sample variance of the score's Jacobian. In the context of causal discovery, even prominent algorithms such as the PC or GES algorithms are proved at the population level and rely on different hyperparameters to work in finite samples, e.g., sparsity coefficient, and alpha level.
Finally, *one way to obtain a principled threshold for our algorithm would require an estimator of the score function with finite-sample guarantees*. Such a result can make a standalone research paper, as the utility of score matching spans various domains, encompassing generative and discriminative models [5,6,7,8,9], and more recently, applications in causal discovery [10] and causal representation learning [11].
*We hope the notes above will help clarify your comments.*
[1]: Thorndike, R. L. (1953). "Who belongs in the family?." Psychometrika.
[2]: V. Satopaa, J. Albrecht, D. Irwin and B. Raghavan. (2011). "Finding a "Kneedle" in a Haystack: Detecting Knee Points in System Behavior." 31st International Conference on Distributed Computing Systems Workshops.
[3]: Goutte, C., Toft, P., Rostrup, E., Nielsen, F. Å., & Hansen, L. K. (1999). "On clustering fMRI time series". NeuroImage.
[4]: Dangeti, P. (2017). Statistics for machine learning. Packt Publishing Ltd.
[5]: Song, Y., & Ermon, S. (2019). Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems, 32.
[6]: Zimmermann, R. S., Schott, L., Song, Y., Dunn, B. A., & Klindt, D. A. (2021). Score-based generative classifiers. arXiv preprint arXiv:2110.00473.
[7]: Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., & Poole, B. (2020). Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456.
[8]: Song, Y., Garg, S., Shi, J., & Ermon, S. (2020, August). Sliced score matching: A scalable approach to density and score estimation. In Uncertainty in Artificial Intelligence (pp. 574-584). PMLR.
[9]: Song, Y., & Ermon, S. (2020). Improved techniques for training score-based generative models. Advances in neural information processing systems, 33, 12438-12448.
[10]: Rolland, P., Cevher, V., Kleindessner, M., Russell, C., Janzing, D., Schölkopf, B., & Locatello, F. (2022, June). Score matching enables causal discovery of nonlinear additive noise models. In International Conference on Machine Learning (pp. 18741-18753). PMLR.
[11]: Varici, B., Acarturk, E., Shanmugam, K., Kumar, A., & Tajer, A. (2023). Score-based causal representation learning with interventions. arXiv preprint arXiv:2301.08230.