Spatial-Aware Multi-Task Learning Based Speech Separation

Online meetings have become an indispensable part of our lives. However, background noise from other family members, roommates, and office mates not only degrades the voice quality but also raises serious privacy issues. In this paper, we develop a novel system, called Spatial Aware Multi-task learning-based Separation (SAMS), to extract audio signals from the target user during teleconferencing. Our solution consists of two novel components: (i) generating fine-grained location embeddings from the user's voice and inaudible tracking sound, which contains the user's position and rich multipath information, (ii) developing a source separation neural network using multitask learning to jointly optimize source separation and location.

Paper

Similar papers

© 2026 NYSGPT2525 LLC