We thank the reviewer for providing feedback regarding use of data in our submission.
The benchmarks for audio-visual data for training and testing of models used in our work are from VGGSound [1] and AudioSet [2]. These are standard open creative commons and are datasets and benchmarks for audio-visual learning tasks as well as audio generation tasks. For example, in the video-to-audio generation literature, prior state of the art works such as Diff-Foley [3], FoleyGen [4], SpecVQGAN [5] and Im2wav [6] use VGGSound for training and testing. Similarly, in audio-visual understanding literature, CAV-MAE [7] utilizes both VGGSound and AudioSet-2M for training and testing on audio-video retrieval tasks. Therefore, it is a standard practice to use these two open and publicly released datasets and to compare novel proposed approaches performance on the same datasets with these prior works in the field.
Similar to prior works, we extracted audio and visual features from these videos in the aforementioned datasets using deep neural networks from publicly released pretrained vision and audio backbones for training and testing our models and other baselines. We do not keep the visual frames and the audio waveforms besides a few samples which are used for demonstration purposes, i.e., qualitative samples to evaluate qualitatively the approach performance. These few samples are not publicly available for demonstration purposes and are provided exclusively via anonymous link (webpage) for reviewers for confidential review. Upon publication of our work we will make sure that the published samples in the manuscript or on the public webpage have explicit permission by creators for the publication.
Besides samples, upon publication of our work we do not intend to publish data besides synthetic text we created for training and testing of our model and the pretrained model for research purposes. We will consult and obtain permission from our institution/s and follow NeurIPS guidelines if this becomes relevant.
We also wanted to note that in our work we performed a user study and included the results of the study in the manuscript. We followed our institutional process for determination whether IRB approval is applicable, and upon determination that it is, we have applied and obtained such approval from our institution. We will follow the reviewer recommendation and state that explicitly in the final version of the manuscript.
Sincerely,
Authors of Paper Submission 1915
References:
[1] Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. Vggsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 721–725. IEEE, 2020.
[2] Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 776–780. IEEE, 2017.
[3] Simian Luo, Chuanhao Yan, Chenxu Hu, and Hang Zhao. Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 48855–48876. Curran Associates, Inc., 2023.
[4] Xinhao Mei, Varun Nagaraja, Gael Le Lan, Zhaoheng Ni, Ernie Chang, Yangyang Shi, and Vikas Chandra. Foleygen: Visually-guided audio generation. arXiv preprint arXiv:2309.10537, 2023.
[5] Vladimir Iashin and Esa Rahtu. Taming visually guided sound generation. arXiv preprint arXiv:2110.08791, 2021.
[6] Roy Sheffer and Yossi Adi. I hear your true colors: Image guided audio generation. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.
[7] Gong, Y., Rouditchenko, A., Liu, A. H., Harwath, D., Karlinsky, L., Kuehne, H., and Glass, J. Contrastive audio-visual masked autoencoder. arXiv preprint arXiv:2210.07839, 2022b.