Further Explanation and Clarification
Thank you once again for your comments on our paper. We appreciate your concerns regarding the effectiveness of the model architecture, and we will provide further clarification on this matter.
---
Firstly, we acknowledge that the proposed dataset can significantly enhance the effectiveness of the model in understanding video anomalies.
**However, it is unreasonable to conclude that the motion model has limited significance based on this.**
Here are some reasons:
---
Firstly, to demonstrate the effectiveness of motion information, we have conducted a rigorous ablation study as in Table 3 and 4 of the main paper. The experimental results, including quantitative metrics and qualitative analysis, indicate a significant improvement when motion information is integrated, compared to **"w/o motion information"**, across **ALL** Text-level and GPT-guided metrics. (**However, Video-ChatGPT and VideoChat only utilize a single type of GPT-guided metric.**)
In the magnitude of performance improvement, integrating motion information resulted in an average performance increase of around **2.4%** in GPT-guided metrics. In comparison, the performance increase in Video-ChatGPT compared to the SOTA baseline was around **1.8%** in their paper. **Therefore, the performance improvement observed in the ablation study is significant.**
**Hence, the unfounded assertion by the reviewer that the improvement in model performance is merely "comparable" is both unreasonable and lacks sufficient evidence.**
---
Secondly, to further demonstrate the effectiveness of motion information, we compared the results of fine-tuning other baselines on the same training data. **This type of experiment serves as a valid form of evidence (acknowledged by Reviewer K7jv and previously used in VideoChat)**.
When compared to other existing baselines, our framework significantly outperformed after fine-tuning on the same training and testing data. Since other baselines do not incorporate motion information, this comparison showcases the advantage of our framework in leveraging motion for video anomaly understanding.
Additionally, it is worth noting that the performance of our base model ("w/o motion information" in Table 3 (A)) was **initially weaker** than Video-LLaMA and LlaMA Adapter in the same training data. **However, after integrating motion-related information and motion-related loss functions, the performance saw a significant enhancement.** This demonstrates that the performance improvement of the model is indeed derived from motion information.
---
Thirdly, the problem that this paper aims to address is to enhance the system's ability to understand anomalous information in videos. Although our method is built upon the framework of video understanding, it would be **unfair** to directly assess the novelty of our approach based on previous general video understanding frameworks. Instead, our improvements are significant in the field of understanding abnormal information in videos. **This contribution has been acknowledged and praised by Reviewers K7jv, Cy7X, and kzaw in the Strengths section.**
---
Fourthly, we have noticed in the supplementary comments that while **the first point is to acknowledge our significant contribution to the dataset** (even considering it the sole contribution), however, **the third point denies the contribution of certain parts of the dataset** (stating that it doesn’t seem to offer substantial advancements over existing datasets). This raises doubts about whether your comments are made from a reasonable and fair perspective. We look forward to further discussion to address and resolve these perplexities.
---
Thanks again for your response and look forward to your response.