Response to Reviewer p9QH (Part 2/2)
**[About Potential Failure Cases]**
- We have indeed discussed potential failure cases in the limitations section (L716-L722) of the appendix. Specifically, when the visual scene becomes highly complex or the video is exceptionally long, synchronization accuracy can be constrained by the performance of the audio signal estimation and the quality of the training data.
- Based on your suggestions, we have also visualized and included some failure cases in the revised version. Please refer to the uploaded revision for our visualization analysis. To address these challenges in temporal alignment, we plan to focus on constructing highly visual-audio-aligned datasets and advancing model design in future work.
**[About Details of Subjective Evaluations]**
Thank you for pointing out this issue. In addition to the details provided in Section B.3 of the appendix, we will incorporate the following clarifications in the final version of the paper, as per your suggestions.
We conducted a user study involving **20 participants** who rated 40 randomly selected video-audio samples, following V2A-Mapper. Specifically, the participants were either practitioners in audio generation and multimedia or PhD students specializing in Artificial Intelligence. To ensure unbiased feedback, all results were presented to the participants anonymously.
Evaluating audio quality, temporal alignment, and semantic alignment across samples from different models requires significant focus from participants. To reduce cognitive load and avoid random ratings, we opted for **pairwise comparisons** instead of Mean Opinion Scores or Meaningful Difference Scores, as used in V2A-Mapper. As shown in Figure 3 of the appendix, pairwise comparisons simplify the evaluation process by asking participants to compare two results at a time, which improves the reliability of their judgments and ensures higher utilization of user study votes. Such a pairwise comparison design is also widely used in the fields of large language models (LLMs) [1,2] and computer vision [3,4]. Specifically, in each trial, participants were presented with two results: one generated by FoleyCrafter and the other by a randomly selected baseline.
To further ensure the quality and consistency of the evaluations, we designed the study to present the same questions multiple times throughout the process. This repetition helped us verify participants’ attentiveness and identify any inconsistencies in their responses. Inconsistent scores for the same question from the same participant were treated as unreliable, allowing us to maintain the integrity of the results.
[1] Liu, Yinhong, et al. "Aligning with human judgement: The role of pairwise preference in large language model evaluators." arXiv preprint arXiv:2403.16950 (2024).
[2] Liusie, Adian, et al. "Efficient LLM Comparative Assessment: a Product of Experts Framework for Pairwise Comparisons." arXiv preprint arXiv:2405.05894 (2024).
[3] Li, Shufan, et al. "Aligning diffusion models by optimizing human utility." arXiv preprint arXiv:2404.04465 (2024).
[4] Zeng, Yanhong, et al. "Aggregated contextual transformations for high-resolution image inpainting." IEEE Transactions on Visualization and Computer Graphics 29.7 (2022): 3266-3280.