A complaint to Reviewer zFyv
Dear AC,
Thanks for your attention to our paper. Here we would like to make a complaint against **reviewer zFyv** for his/her extreme irresponsible review. We argue his/her professionalism from the following three perspectives.
(1)First, **limited paper reading capability**. He/her reviewed in weakness1 that “paper is not easy to follow”. We think this review is not appropriate due to that **two other reviewers have explicitly reviewed that our paper is easy to follow**. For example, reviewer 1sdz has reviewed that “**The paper is clearly written and easy to read**” and reviewer hHHx has reviewed that “**the paper is well-written with clear explanations**.” So we think probably reviewer zFyv himself/herself dose not have enough paper reading capability and patience.
(2)Second, **absurd judgement about “incremental”**. He/Her reviewed in weakness2 and weakness4 that ”Paper is highly incremental”, for the reason that “model have been heavily built on top of numerous existing pertained models” or ”VAST stands on the shoulders of numerous pertained models developed by others”.
+ The whole VAST model or trained vision/audio captioners only takes **three** current models (BERT for text, CLIP for vision and BEATs for audio) for single-modality encoder initialization, which is exaggerated by him/her with words **“heavily”** and **“numerous”**. In addition, those three models we take are very common and used by lots of cross-modality researchers. As stated in our rebuttal to him/her for weakness2, cross-modality pretraining which is the theme of our paper concentrate more on multi-modal connections training instead of single-modality representation training. Using those weights as initialization can be a good start point letting model focus on cross-modality learning, which is also high-efficiency especially for researchers with limited computation source.
+ According to his/her standard to infer if a work is “incremental”, he/her can **reject any paper** even without reading rather than those 1) who happenedly researches image classification, action recognition, audio recognition. 2) who has enough computation source to train all things from scratch (such as CLIP, CoCa). So the review process will be extreme easy that we get a paper, and we directly look into the implement details, and if we successfully find that current work does not belong to above two classes, we just tag it as “incremental” and give a reject score.
+ The contribution of VAST modal and VAST-27M dataset has been illustrated in main paper as well as in [common question 1] in our global rebuttal, and also **recognized by all other reviewers**. For example, reviewer v49r reviewed that ”The new VAST-27M dataset could be very valuable for researchers working on related topics. The experimental results also demonstrate the advantages of the trained model“. Reviewer hHHX reviewed that ”I recognize the contributions of the paper, especially the dataset”. Reviewer “1sdZ” reviewed that "The idea of combining text, video, and audio in one caption is interesting and novel. I believe this will contribute to the multimodal learning community.“ In addition, all reviewer think that the performances of VAST on series of benchmarks are strong.
(3) Third, **unfamiliar with cross-modality pretraining’s common objectives**. Two questions (Q2 and Q5) proposed by him/her concentrate on contrastive and matching loss, which are two extreme common losses widely used in cross-modality pretraining field for relatively a long time. VAST uses them to strengthen both coarse-grained and fine-grained alignment. In Q2, he asks the necessity of matching loss and in q5 he asks what if ratio of two losses are 1:1, however, the ratio of two losses are just 1:1 as shown in Eq4 of the main paper. Those questions reflect that he/her **has not read paper carefully** and is also **not familiar with the basic knowledge of cross-modality pretraining**.
(4) Fourth, **invaluable feedback**. We sincerely thanks for NIPS that support a discussion period that let authors and reviewers communicate to each other. The authors have **carefully answered his all questions and shown many additional experiment results in the rebuttal stages**. However, as you commented, he/she commented with simply “maintain my original rating” without any reason or further questions.
In a word, we think it is not appropriate to give a 3 score reject rating to our paper with such subjective reasons (individually fell not easy to read / individually feel incremental), and we think reviewer zFyv is not professional enough to serve as a reviewer. We sincerely hope AC could judge this paper and reviewer zFyv ’s review fairly, thanks very much!
Best,
Author