> While the authors provide a general method for learning equivariant, it would be beneficial to compare it with a hardwired equivariant to show its competitive performance.
We want to stress that we make no claims that our method for learning an equivariant f_w will match the performance of a hardwired f_w, in fact, we acknowledge in the paper that having some equivariances hardwired would likely lead to an increase in performance.
However, we agree that comparison with a hardwired equivariant $f_\omega$ would be interesting. If you have suggestions or know of relevant work for how to parameterise such an $f_\omega$, we'd be happy to run those experiments. This would have been easy if $f_\omega$ was a function from image to image space equivariant to translations/rotations, but we don't know of any work that does this for functions from an image space to transformation-parameter space.
> I remain unconvinced that boundary effects are insignificant. The method’s core idea, when used as a preprocessing step for VAE, is to augment samples in the group orbit accurately. Although GalaxyMNIST does not vanish at the boundary, it primarily consists of sparse objects on a black background.
As we mentioned in our previous response, we acknowledge that this is a potential limitation of our method, but note that this is also a limitation of existing published methods. Thus, we hope that we will not be held to a higher standard for publication.
> The paper assumes prior knowledge of the group symmetry in the dataset. If I understand correctly, the comparison to [Yang et al. 2023] might be unfair, as [Yang et al. 2023] genuinely learns the general linear group, whereas SGM assumes the transformation group is limited to the rotation and scaling subgroup.
To clarify, in our experiments for this paper our SGM covers rotation, scaling, *and shifting*. This makes it slightly less flexible than the LieGAN of Yang et al. [2023], since their method is also able to learn shearing and flipping. However, we feel that the comparison is largely fair, since we are not reporting any quantitative results, and instead focus on qualitative comparisons. From these qualitative comparisons, it is clear that fundamental differences between our two approaches (e.g., our SGM learning conditional distributions) are the dominant reason for different behavior, rather than the choice of assumed transformation group. We also note that our SGM is also capable of learning shearing and flipping, we simply didn't include these transformations in our experimental results. However, expending our assumed transformation group to include these is trivial.
> Nonetheless, since my main concern about demonstrating the correct conditional distribution learning (points 1, 2, and 6 in my original review) has been addressed, I will substantially raise the score while still reserving my opinion on other aspects.
Thank you for increasing your score and engaging with the rebuttal. We very much appreciate the discussion and we are glad that the rebuttal has already addressed many of your concerns. We hope that we have addressed your remaining concerns and that you will consider increasing it further.