Thanks for your feedback
Thank you for your insightful comments and for adjusting the scores. We would like to clarify a few points where we believe there might have been a misinterpretation of our work.
For condition (ii), we already provide its theoretical interpretability in Corollary1 and App. A.5. This condition is a straightforward lower bound for identifiable causal effect, which demonstrates **for the first time** how strong causal effects can be recognized. For practical testability of condition (ii), as given in Q&A 2.3 of Reviewer 8ao9, **most of conditions is unverifiable without the ground truth SEM (even simple conditions like linear and nonlinear)**. Therefore, to further provide more practical testability, additional experiments are given on synthetic data. We test different setting of the sum of causal effect to show the performance of CaPS when condition (ii) are perfectly satisfied / likely satisfied / likely unsatisfied. In order to accurately control the causal effect and the lower bound in condition (ii), we use the linear SynER1 with the noise standard deviation in $U(0.4, 0.8)$. Under this settings, condition (ii) will perfectly satisfied when the minimal causal effect is geater than $\sqrt{0.8^2(\frac{1}{0.4^2}-\frac{1}{0.8^2}))}=\sqrt{3}$ because the node with weakest SATE and single child will greater than the theoretical lower bround. The experimental results are shown in Table 5, which gives the practical testability and shows that **CaPS will work well when condition (ii) are perfectly satisfied and likely satisfied**.
**Table 5. practical testability of condition (ii) on SynER1**
| condition (ii) | causal effect | SHD | SID | F1 |
|---------------------|---------------|---------|-----------|--------------|
| perfectly satisfied | $U(1.8, 2.0)$ | 0.0±0.0 | 0.0±0.0 | 1.000±0.000 |
| likely satisfied | $U(1.6, 1.8)$ | 0.0±0.0 | 0.0±0.0 | 1.000±0.000 |
| likely satisfied | $U(1.4, 1.6)$ | 0.2±0.4 | 0.8±1.6 | 0.975±0.050 |
| likely satisfied | $U(1.2, 1.4)$ | 0.4±0.4 | 1.4±1.7 | 0.964±0.049 |
| likely unsatisfied | $U(0.6, 0.8)$ | 2.6±1.0 | 10.0±5.6 | 0.772±0.065 |
| likely unsatisfied | $U(0.4, 0.6)$ | 4.2±1.4 | 8.8±6.8 | 0.689±0.082 |
| likely unsatisfied | $U(0.2, 0.4)$ | 4.0±1.0 | 15.8±11.7 | 0.655±0.103 |
For zero-mean Gaussian, this settings is widely used in previous work (ref. [9,11,13,19,...]) because the $\epsilon$ are usually consider as the residual of $f(pa(x))$. As LISTEN have pointed out, "without loss of generality, we assume that $E(X_i)=E(N_i)=0$". This is because the non-zero-mean and zero-mean is equivalent in a ANM, which we already explained in Q&A 1.2. Thus, any non-zero-mean ANM with $\epsilon_n \sim N(\mu,\sigma)$ can transform to a zero-mean ANM with $\epsilon_z \sim N(0,\sigma)$ then follow the same derivation. And that's why $\mu$ does not significantly affect the relative performance empirically in Table 1 & 2.
For Gaussian noise assumption, to be precise, we use Gaussian pdf for derive Eq.2 and Eq.10 comes from Eq.2. Since CaPS **handles linear, nonlinear and even mixed relations simultaneously**, we put almost no restrictions on $f$ in ANM. Under this premise, as CaPS handles both types of relations for the first time, it seems too strict to require a solution for both linear & nonlinear and Gaussian & non-Gaussian at the same time. Although this is the weakest condition we can derive for both linear and nonlinear scenario, the experimental results in App. C.7 have been encouraging in cases of non-Gaussian noise. We are striving to broaden it to more relaxed conditions.
Once again, thanks for your careful review and the time you have dedicated to our manuscript. We hope that our responses will address your concerns.