Deep Feature CycleGANs: Speaker Identity Preserving Non-parallel\n Microphone-Telephone Domain Adaptation for Speaker Verification
With the increase in the availability of speech from varied domains, it is\nimperative to use such out-of-domain data to improve existing speech systems.\nDomain adaptation is a prominent pre-processing approach for this. We\ninvestigate it for adapt microphone speech to the telephone domain.\nSpecifically, we explore CycleGAN-based unpaired translation of microphone data\nto improve the x-vector/speaker embedding network for Telephony Speaker\nVerification. We first demonstrate the efficacy of this on real challenging\ndata and then, to improve further, we modify the CycleGAN formulation to make\nthe adaptation task-specific. We modify CycleGAN's identity loss,\ncycle-consistency loss, and adversarial loss to operate in the deep feature\nspace. Deep features of a signal are extracted from an auxiliary (speaker\nembedding) network and, hence, preserves speaker identity. Our 3D\nconvolution-based Deep Feature Discriminators (DFD) show relative improvements\nof 5-10% in terms of equal error rate. To dive deeper, we study a challenging\nscenario of pooling (adapted) microphone and telephone data with data\naugmentations and telephone codecs. Finally, we highlight the sensitivity of\nCycleGAN hyper-parameters and introduce a parameter called probability of\nadaptation.\n