We sincerely thank Reviewer T5M6 for the thoughtful feedback and valuable suggestions.
---
## T5M6-Q1. Quality control of the dataset
We established the reliability of our dataset's quality in two aspects. First, the quality of audio generated using YT-Ambigen is in line with well-established large-scale benchmarks, e.g., VGGSound. As reported in Table 2, the FAD (3.95 vs. 3.62) and KLD (1.77 vs. 2.23) scores of our model trained with YT-Ambigen's mono channel are similar to those of VGGSound. These metrics from our VGGSound-trained model are also comparable to the prior arts like SpecVQGAN and Diff-Foley, suggesting that YT-Ambigen offers a reliable data source for benchmarking video-to-audio generation.
Second, using a combination of performant off-the-shelf multimodal discriminators for filtering is a well-established practice for constructing large-scale datasets with quality control [1, 2, 3], and could be more effective than manual filtering in some cases. For instance, in Table 2, our predecessor in discriminative audio-visual reasoning (YT360) reports an FAD metric of 15.91 for video-to-audio generation, despite being collected through manual filtering.
---
## T5M6-Q2. Distribution statistics of the dataset
Thank you for your suggestion. We included the distribution statistics of YT-Ambigen in the Appendix G, which cover:
- (a) The top-50 AudioSet label distribution predicted with PaSST [1]
- (b) The top-50 COCO object class distribution of the most salient object per video with FPN [2]
- (c) The center coordinates of each salient object’s bounding box
- (d) The tracking of center pixels per video predicted with CoTracker [3] (randomly selected 1K samples for visibility).
Our audio distribution is similar to that of AudioSet, where YT-Ambigen covers 314 out of 527 classes in AudioSet, accounting for 97.91% of the entire AudioSet videos. Moreover, the semantic, spatial, and temporal distributions of the most salient object per video are summarized in (b-d). These objects cover 79 out of 80 classes in COCO. Moreover, they are located in diverse positions within the field of view and often move around during five-second segments, creating more challenging scenarios for video-to-ambisonics generation.
---
## T5M6-Q3. More examples covering challenging scenarios
We included more qualitative examples with non-centered or moving objects in our demo page and Appendix F.
---
## T5M6-Q4. Codebook generation with nine residual codebooks
We apologize for the confusion. As shown in the legend of Figure 2, each block represents a codebook group rather than an individual code. All residual codes from the selected groups are generated at each sequence step. For example, with 9 RVQ codes per channel:
- Step 1: 1 code from $W_p$
- Step 2: 8 codes from $W_r$ and 3x1 from $S_p$ (total: 11)
- Step 3: 1 code from $W_p$ and 3x8 from $S_r$ (total: 25)
- and so on.
The figure caption has been updated to clarify this process.
---
## T5M6-Q5. Typos and visualization suggestions
Thank you for your suggestions. We fixed the typos and updated visualizations in Figure 3 and 5.
---
[1] Lee et al. Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation Learning. In ICCV 2021.
[2] Nagrani et al. Learning Audio-Video Modalities from Image Captions. In ECCV 2022.
[3] Wang et al. A Large-scale Video-Text Dataset for Multimodal Understanding and Generation. In ICLR 2024.