Further Clarifications
We agree with the reviewer that we have not proven other approaches will not work. Indeed, let us consider the following argument. We'll work with $H$ a Hilbert space of real-valued functions over the domain $D = [0,1]^2$. Consider first the following definition of a video: A two-frame video is a pair of functions $f_0, f_1 \in H$ such that there exists a bounded, injective map $T: D \to \mathbb{R}$ with the property that $f_1(x) = f_0 \big ( T^{-1}(x) \big )$ for all $x \in T(D) \cap D$.
Now let $G: H \to H$ be a generative model for which there exists some $\xi_0 \in H$ such that $G(\xi_0) = f_0$. Then following holds: if $G$ is equivariant w.r.t. to $T$ then the pair $\big ( f_0, G(\xi_0 \circ T^{-1}) \big )$ is a two-frame video. This follows from the definition of equivariance as $G(\xi_0 \circ T^{-1}) = G(\xi) \circ T^{-1} = f_0 \circ T^{-1} $ on $T(D) \cap D$. Certainly, the other direction does not hold true. In particular, it does not follow that if $\big ( f_0, G(\xi_0 \circ T^{-1}) \big )$ is a two-frame video then $G$ is equivariant w.r.t. $T$. Therefore we agree with the reviewer that we have not proven that equivariance is necessary, however, we have proven that it is sufficient. We will explicitly point this out in the camera-ready version.
This lack of necessity stems from the fact that our definition of a video is weak and allows many possible pairs to be videos. One natural way of strengthening it is as follows: Let $\mathcal{T}$ denote the set of bounded, injective maps on $D$ then a two-frame video is a pair of functions $f_0,f_1 \in H$ such that the following problem admits a unique maximizer
\[\max_{T \in \mathcal{T}} \big \{ |T(D) \cap D| : f_1(x) = f_0 \big ( T^{-1}(x) \big ) \text{ for all } x \in T(D) \cap D \big \}.\]
In particular, we enforce that there is a unique deformation which keeps the maximum number of pixels in frame. From uniqueness, it now follows that: the pair $\big ( f_0, G(\xi_0 \circ T^{-1}) \big )$ is a two-frame video if and only if $G$ is equivariant w.r.t. to $T$.
We thank the reviewer for bringing this to our attention and are happy to discuss the mathematical modeling of videos in the camera-ready version further.
If that's not what the Reviewer was looking for, please let us know and we will try our best to incorporate the Reviewer's feedback in our camera-ready.