Response to the Reviewer
We thank the reviewer for the careful review and insightful comments. Below we address the questions and comments raised in this review.
* W1: Please find our detailed response to the questions.
* W2: Thank you for pointing out the issue with notation overlap. We have changed $a = x + iy$ into $a = v + iw$, where we refrained from using $a = u + iv$ since $\mathbf{u}$ is reserved for the inputs of the LTI systems.
* W3: We have added the integral definition of the total variation to our manuscript.
* Q1: The reviewer has correctly identified that a high-magnitude transfer function is sufficient for a frequency to "pass." Here, we provide more intuitions in terms of why the total variation is a good measurement of the frequency bias. As noted in [2], one reason why an LTI system in an SSM degenerates is that its transfer function is too flat. That is, if $\mathbf{G}$ can be well-approximated by a much lower-degree rational function, then one can essentially replace the large LTI system with a much smaller one. In other words, our LTI system is not as expressive as its size may have suggested. If we now restrict our attention to only a part of the frequency domain, then the same reasoning applies: if the transfer function $\mathbf{G}$ is flat on $[a,b]$, then we can replace a large LTI system with a much smaller one. While the small system may not well-approximate the original system on the entire frequency domain, the approximation is good on $[a,b]$. Hence, this essentially means that our LTI system is unable to capture "complex patterns" in the frequency interval $[a,b]$ because all it does in $[a,b]$ can also be done by a much smaller system. We have added a similar discussion of this in Appendix A.
* Q2: As $L \rightarrow \infty$, $\tan((1 - (L-1)/L) \pi / 2) \rightarrow \infty$ so the "edge of nonzero information" could grow arbitrarily large. However, this is only due to the fact that very few samplers of the transfer function $\mathbf{G}$ will become large. See Figure 7 (Right), for example. While there exists a sampler whose imaginary part $> 10000$, most are actually $< 1000$. That being said, as $L \rightarrow \infty$, one can always make up a synthetic input data of which $\alpha$ needs to be arbitrarily large to capture. However, in the average case, the few large samplers would only contain a vanishing proportion of the information as $L \rightarrow \infty$ --- that is why Rule I does not involve $L$.
* Q3: The reviewer raised a good point. Different choices are made in the definition of the Sobolev norm. These norms are different but norm-equivalent (i.e., $c \\|\cdot\\|_x \leq \\|\cdot\\|_y \leq C \\|\cdot\\|_x$ for constants $c, C > 0$) so that they lead to the same Sobolev space. In our case, the $(1+|s|^2)^\beta$ and $(1+|s|)^{2\beta}$ factors lead to equivalent norms, although they are numerically different. The factor we chose came from the spherical harmonics community [1], but we have remarked that in our manuscript to avoid potential confusion.
* Q4: We have added a concrete guideline for tuning $\alpha$ in Appendix G. For the LRA tasks, we do not set $\beta$ as a tuning parameter but instead make it a trainable parameter. In general, we find that on LRA tasks, tuning $\alpha$ is indeed more effective than training $\beta$. However, our ablation study below still shows that the model benefits from having a $\beta$ involved.
| | ListOps | Text | Retrieval | Image | PathFinder | PathX |
|:-------:|:-------:|:-------:|:-------:|:-------:|:-------:|:-------:|
|only $\alpha$|61.24|89.32|91.35|90.51|95.74|97.52|
|only $\beta$|61.15|88.92|90.31|89.75|95.93|96.12|
|with neither|60.47|86.18|89.46|88.19|93.06|91.95|
|with both|62.75|89.76|92.45|90.89|95.89|97.84|
Moreover, for some tasks where one needs a more extreme frequency bias, having only $\alpha$ is suboptimal, while $\beta$ is really making a big difference. (See Table 2 for example.)
We hope this answers the reviewer's questions and concerns. We are happy to answer any follow-up question(s) that the reviewer may have.
[1] J. A. Barcelo, M. Folch-Gabayet, T. Luque, S Perez-Esteva, and M. C. Vilela. The Fourier extension operator of distributions in Sobolev spaces of the sphere and the Helmholtz equation. Proc. Roy. Soc. Edin. Sec. A: Math., 151(6):1768–1789, 2021.
[2] A. Yu, M. W. Mahoney, N. B. Erichson, HOPE for a robust parameterization of long-memory state space models, arXiv preprint arXiv:2405.13975, 2024.