FakeMyVoice: A Hybrid Multi-Stage Deep Learning Pipeline for Low-Resource Speaker-Conditioned Voice Cloning with Adversarial Robustness and Ethical Safeguards

Voice cloning has rapidly advanced from a research curiosity to a practical capability with far-reaching applications in accessibility, entertainment, and human-computer interaction. However, existing systems typically demand substantial reference audio (often 30+ minutes) or rely on end-to-end architectures that obscure internal representations and resist interpretability. This paper presents FakeMyVoice, a modular, interpretable, and low-resource voice cloning framework that integrates a Generalized End-to-End (GE2E) speaker encoder, a speaker-conditioned Tacotron 2 synthesizer, and a WaveGlow flow-based neural vocoder into a coherent three-stage pipeline. Our system achieves a speaker verification accuracy of 92% across 40+ speakers using compact 256-dimensional embeddings and produces intelligible, speaker-consistent speech from as few as 30 seconds of reference audio. We train on three widely used corpora - LJSpeech, VCTK, and VoxCeleb - and evaluate using Mean Opinion Score (MOS), Speaker Similarity Score (SpeakerSim), Word Error Rate (WER), and Perceptual Evaluation of Speech Quality (PESQ). FakeMyVoice demonstrates competitive performance against state-of-the-art baselines while maintaining full modularity: each component can be independently retrained or replaced. We further propose novel experimental axes including low-resource ablations (30s–5min reference audio), cross-lingual generalization to Hindi and Marathi, and integration with automatic deepfake detection, establishing a complete responsible AI framework around voice synthesis technology.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC