Adversarial Disentanglement of Speaker Representation for Attribute-Driven Privacy Preservation
In speech technologies, speaker's voice representation is used in many\napplications such as speech recognition, voice conversion, speech synthesis\nand, obviously, user authentication. Modern vocal representations of the\nspeaker are based on neural embeddings. In addition to the targeted\ninformation, these representations usually contain sensitive information about\nthe speaker, like the age, sex, physical state, education level or ethnicity.\nIn order to allow the user to choose which information to protect, we introduce\nin this paper the concept of attribute-driven privacy preservation in speaker\nvoice representation. It allows a person to hide one or more personal aspects\nto a potential malicious interceptor and to the application provider. As a\nfirst solution to this concept, we propose to use an adversarial autoencoding\nmethod that disentangles in the voice representation a given speaker attribute\nthus allowing its concealment. We focus here on the sex attribute for an\nAutomatic Speaker Verification (ASV) task. Experiments carried out using the\nVoxCeleb datasets have shown that the proposed method enables the concealment\nof this attribute while preserving ASV ability.\n