From Inference to Generation: End-to-end Fully Self-supervised Generation of Human Face from Speech
This work seeks the possibility of generating the human face from voice\nsolely based on the audio-visual data without any human-labeled annotations. To\nthis end, we propose a multi-modal learning framework that links the inference\nstage and generation stage. First, the inference networks are trained to match\nthe speaker identity between the two different modalities. Then the trained\ninference networks cooperate with the generation network by giving conditional\ninformation about the voice. The proposed method exploits the recent\ndevelopment of GANs techniques and generates the human face directly from the\nspeech waveform making our system fully end-to-end. We analyze the extent to\nwhich the network can naturally disentangle two latent factors that contribute\nto the generation of a face image - one that comes directly from a speech\nsignal and the other that is not related to it - and explore whether the\nnetwork can learn to generate natural human face image distribution by modeling\nthese factors. Experimental results show that the proposed network can not only\nmatch the relationship between the human face and speech, but can also generate\nthe high-quality human face sample conditioned on its speech. Finally, the\ncorrelation between the generated face and the corresponding speech is\nquantitatively measured to analyze the relationship between the two modalities.\n