Thank you for your constructive comments, and for clarifying that you are not opposed to accepting our work. We agree that generalizing our current data distribution would greatly strengthen our results. After a careful examination of our proof, we believe that the data assumptions on $y_i$ in Assumption 1 could be extended to the following form:
(1) $\mathbb{E}(y_i y_j | \mathbf{x}\_{1:N}] = 0$ and $\mathbb{E}[y_i^2 | \mathbf{x}\_{1:N}] = 1$ for all $i \neq j, i,j \in[N]$.
(2) $\mathbb{P}(y\_{1:N} | \mathbf{x}\_{1:N}) = \mathbb{P}(y\_{1:N} | -\mathbf{x}\_{1:N}) $.
Note that such an assumption holds for a wide range of data distributions beyond the case where $\mathbf{x}_i$ and $y_i$ are independent. For example, the following data generating process gives $y_i$ that depends on $\mathbf{x}_i$, but the conditions (1) and (2) above still hold:
Consider an arbitrary fixed vector $\mathbf{a} \in R^d$ with $\|a\|_2 >2$. Suppose that $\mathbf{x}_i$, $i=1,\ldots,N$ are independently generated from the uniform distribution on the unit sphere, and supposed that given $\mathbf{x}_i$, $y_i$ is generated as follows:
- $y_i = 0 $ with probability $1- \frac{1}{ \max \\{\langle \mathbf{a}, \mathbf{x}_i \rangle^2, 1\\}} $,
- $y_i = \max \\{|\langle \mathbf{a}, \mathbf{x}_i \rangle|, 1\\}$ with probability $\frac{1}{2 \max \\{\langle \mathbf{a}, \mathbf{x}_i \rangle^2, 1\\}} $,
- $y_i = -\max \\{|\langle \mathbf{a}, \mathbf{x}_i \rangle|, 1\\}$ with probability $\frac{1}{2 \max \\{\langle \mathbf{a}, \mathbf{x}_i \rangle^2, 1\\}} $.
It is easy to verify that $\mathbb{E}[y_i^2|\mathbf{x}\_{1:N}] = \mathbb{E}[y_i^2|\mathbf{x}_i] = 1$, $\mathbb{E}[y_i y_j|\mathbf{x}\_{1:N}] = \mathbb{E}[y_i y_j|\mathbf{x}\_{i},\mathbf{x}\_{j}] = 0$ and $\mathbb{P}(y\_{1:N} | \mathbf{x}\_{1:N}) = \mathbb{P}(y\_{1:N} | -\mathbf{x}\_{1:N}) $. Moreover, $\mathbf{x}_i$ and $y_i$ are not independent, since $\mathbb{E}[y_i^4| \mathbf{x}_i ] = \max\\{\langle \mathbf{a}, \mathbf{x}_i \rangle^2, 1 \\}$ is a function of $\mathbf{x}_i$.
We will update the paper to include the more general setting under conditions (1),(2) above. We assure that such an extension only requires minor modifications in the paper, and the proofs do not need any significant change. We believe that including such an extension can significantly strengthen our paper, and we hope that it can address your concerns on the limitation of our data models.