Thanks for your constructive suggestion
**Experiment setting.** Thanks for the constructive question and positive response. Your suggestion serves to enhance our paper's clarity, and we welcome the opportunity to provide a comprehensive response. Per your suggestion, we present a comprehensive comparison between the cross-attention maps produced by the DDIM inversion technique and the proposed method within the context of real-image editing. We select a representative image featuring a white cat$\rightarrow$dog, located in Figure 13, row 5. Subsequently, we embark on the visualization of cross-attention maps, focusing on a designated ```dog``` token, during the image generation with (**a**) DDIM inversion or (**b**) the proposed method. In consonance with your suggestion, we chart the evolutionary trajectory of cross-attention maps through sequential sampling steps, each spaced at intervals of 10 steps (0 steps, 10 steps, 20 steps, ...).
**Results.** Upon the initial phase (0 steps), no substantial distinction is evident between the attention maps, as the attention values in both maps are quite randomly distributed. This similarity arises from the initial absence of semantic cues related to the depicted dog within the spatial query.
However, as the process unfolds, marked differences emerge. Specifically, our attention map progressively converges upon key discriminative features distinguishing between cat and dog, such as the face, ears, etc, with minimal emphasis on the structural attributes of the source image, such as background and posture. In contrast, the DDIM inversion's attention map undergoes a distinct transformation. It progressively shifts focus away from the original source image's overarching structure, and instead converges on the newly generated dog subject. This shift entails modifications in positioning, posture, identity, or gaze orientation compared to the source image. We note that these differences between attention maps begin to emerge at early steps, i.e. 10 steps.
These findings suggest that the proposed method effectively maintains the overarching structure of spatial query representations while also adaptively reflecting the semantically encoded contextual information to the designated query representations. It is noteworthy that the control of step size ($\gamma$) and a residual path for context update (Eq. (12)) may avoid the catastrophic devastation of the fundamental structure of the source image. While we're unfortunately constrained from uploading supplementary figures or links at this moment, we will certainly incorporate these new results into the revised paper. Again, we extend our thanks to the reviewer for the careful and constructive suggestion.