Leveraging speaker attribute information using multi task learning for speaker verification and diarization

Deep speaker embeddings have become the leading method for encoding speaker\nidentity in speaker recognition tasks. The embedding space should ideally\ncapture the variations between all possible speakers, encoding the multiple\nacoustic aspects that make up a speaker's identity, whilst being robust to\nnon-speaker acoustic variation. Deep speaker embeddings are normally trained\ndiscriminatively, predicting speaker identity labels on the training data. We\nhypothesise that additionally predicting speaker-related auxiliary variables --\nsuch as age and nationality -- may yield representations that are better able\nto generalise to unseen speakers. We propose a framework for making use of\nauxiliary label information, even when it is only available for speech corpora\nmismatched to the target application. On a test set of US Supreme Court\nrecordings, we show that by leveraging two additional forms of speaker\nattribute information derived respectively from the matched training data, and\nVoxCeleb corpus, we improve the performance of our deep speaker embeddings for\nboth verification and diarization tasks, achieving a relative improvement of\n26.2% in DER and 6.7% in EER compared to baselines using speaker labels only.\nThis improvement is obtained despite the auxiliary labels having been scraped\nfrom the web and being potentially noisy.\n

Paper

References (29)

Scroll for more · 17 remaining

Similar papers

© 2026 NYSGPT2525 LLC