Summary
* The paper focuses on zero-shot voice conversion and presents a framework for training a voice conversion model. The authors introduce a method that leverages entangled speech representations obtained from self-supervised learning (SSL) and speaker verification models. Additionally, they propose a novel training strategy that enhances the synthesis model for voice conversion through the use of self-synthesized examples.
Strengths
* The paper combines strategies from voice conversion and singing voice conversion to improve zero-shot voice conversion. One noteworthy contribution is the introduction of a novel training strategy that utilizes self-synthesized examples for data augmentation and iterative improvement of the generation model. This represents a significant advancement from traditional approaches that heavily relied on heuristic transformations.
* The paper conducts comprehensive experiments. By extensively comparing the proposed framework with several baseline models, the authors effectively demonstrate its efficacy.
* Overall, the paper successfully presents a valuable contribution to the field of voice conversion and offers an approach for achieving zero-shot voice conversion.
Weaknesses
* The main framework presented in this paper appears to draw from existing voice conversion and singing voice conversion techniques, such as utilizing SSL features, speaker embeddings, and incorporating prosody information like duration and pitch. These strategies resemble prior work ([1-3]) in the field. Additionally, the synthesizer's approach showcases similarities to methods used in speech synthesis, specifically resembling techniques employed in FastSpeech 2 [4]. To support these statements and provide a comprehensive overview of the related work in the field, it would be beneficial for the authors to provide appropriate citations and conduct comparative analyses.
* Moreover, the claim of the framework's efficiency in scaling to other languages through the introduction of SSL features is not unique to the authors' proposed model since similar approaches have been explored elsewhere.
* Additionally, it is worth noting that the concept of data augmentation using self-synthesized examples has been discussed in the literature, particularly in the context of speaker verification [5].
* Furthermore, the experimental results suggest that the inclusion of self-synthesized examples for data augmentation yields only a marginal improvement in the model's performance.
* Consequently, a deeper analysis or controlled study could be conducted to better isolate and understand the specific effects of the proposed training strategy from other contributing factors.
References:
[1] Jayashankar T, Wu J, Sari L, et al. Self-Supervised Representations for Singing Voice Conversion[C]//ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5.
[2] Hussain S, Neekhara P, Huang J, et al. ACE-VC: Adaptive and Controllable Voice Conversion Using Explicitly Disentangled Self-Supervised Speech Representations[C]//ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5.
[3] Maimon G, Adi Y. Speaking style conversion with discrete self-supervised units[C]//EMNLP 2023. EMNLP, 2023.
[4] Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,”[C]//International Conference on Learning Representations (ICLR), 2021.
[5] Cai D, Cai Z, Li M. Identifying Source Speakers for Voice Conversion Based Spoofing Attacks on Speaker Verification Systems[C]//ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023: 1-5.
Questions
* The authors should clarify what specific aspects of their framework, beyond the use of SSL features, contribute to this scalability to strengthen the claim of originality in this regard.
* Could the authors elucidate the unique characteristics or modifications they have made within their framework that differentiate it from other SSL-based voice conversion models?
* The introduction of the self-synthesized examples as a data augmentation method is intriguing. However, its efficacy and broader applicability remain questions. Can the authors provide experimental evidence or analysis showcasing the generalizability of this data augmentation technique across various methods?
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.