Thanks for your response
Thank you for your helpful response.
I acknowledge that this work first proposes the generalized model for multiple speech tasks. Moreover, I think your model will have a good impact on speech research. That is why the authors should conduct more experiments fairly for future researches who might follow your research. I don't want anyone in the future to claim "VoiceBox experimented like this and we're just following VoiceBox".
I still have a concern about the comparisons for each task. I hope that you do not cut corners by just arguing that our framework is novel.
>**Zero-shot TTS experiment**
I still disagree with your experiments because YourTTS is not a good zero-shot TTS model. The audio quality is not good and the models are trained with different datasets. It is not fair to compare it with your model. The authors also utilize an additional Duration model.
>**Diverse speech generation experiment**
you just compared the model with VITS-LJ and VITS-VCTK. I could not agree that your model is better than others with the results of this experiment. You should train the model with the same dataset for a fair comparison.
>**To generalize speech infilling, any powerful non-autoregressive generative models, including diffusion models, should work. We chose flow-matching (FM) with optimal transport (OT) path because [1] showed that FM w/ OT > FM w/ diffusion > score-matching (SM) w/ diffusion (the typical diffusion model) on training speed and inference compute-quality trade-off. See the comparison in Table 1 and Fig 4-7 in [1].**
If you're saying this, adopting the flow-matching (FM) with optimal transport (OT) path is not your contribution. the authors should have compared all scenarios (FM w/ OT, FM w/ diffusion, score-matching (SM) w/ diffusion) to verify that the FM w/ OT is also a better method for speech generative tasks.
> **Duration Modeling**
I also have an additional doubt on the comparison of different models. In Appendix, the WER results of flow-matching and regression model are almost same in Table B3. This results show that the trained model with your dataset has a just lower WER so you should train the VITS or other models with the same dataset you used. I think a duration modeling with a large-scale dataset improves the pronunciation. For a fair comparison, VITS with a duration modeling should be compared. Specifically, VITS just utilizes a MAS for a efficient training without external duration modeling. In addition, there are many works which utilizes VITS with external duration modeling to improve the performance.
[1] Cite as: Ju, Y., Kim, I., Yang, H., Kim, J.-H., Kim, B., Maiti, S., Watanabe, S. (2022) TriniTTS: Pitch-controllable End-to-end TTS without External Aligner. Proc. Interspeech 2022, 16-20, doi: 10.21437/Interspeech.2022-925
[2] Cite as: Lim, D., Jung, S., Kim, E. (2022) JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech. Proc. Interspeech 2022, 21-25, doi: 10.21437/Interspeech.2022-10294
[3] Zhang, Yongmao, et al. "Visinger: Variational inference with adversarial learning for end-to-end singing voice synthesis." ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022.
[4] Y. Shirahata, R. Yamamoto, E. Song, R. Terashima, J. -M. Kim and K. Tachibana, "Period VITS: Variational Inference with Explicit Pitch Modeling for End-To-End Emotional Speech Synthesis," ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 2023, pp. 1-5, doi: 10.1109/ICASSP49357.2023.10096480.
Basically, I like the concept of this paper. However, current manuscript does not conducted a fair comparison. I encourage the authors to add additional ablation studies to the paper. But, they did not conduct any experiments I suggested so I could not take any action in this stage.