MULTI-CHANNEL SPEECH SEPARATION

Patent №

US 10,839,822

Granted

2020-11-17

Filed 2017

Owner

MICROSOFT TECHNOLOGY LICENSING, LLC

AI components

4

ml · nlp · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

15805106

Representative embodiments disclose mechanisms to separate and recognize multiple audio sources (e.g., picking out individual speakers) in an environment where they overlap and interfere with each other. The architecture uses a microphone array to spatially separate out the audio signals. The spatially filtered signals are then input into a plurality of separators, so each signal is input into a corresponding signal. The separators use neural networks to separate out audio sources. The separators typically produce multiple output signals for the single input signals. A post selection processor then assesses the separator outputs to pick the signals with the highest quality output. These signals can be used in a variety of systems such as speech recognition, meeting transcription and enhancement, hearing aids, music information retrieval, speech enhancement and so forth.

Machine learningNatural languageSpeechAI hardwareG10L 21/0216G06N 3/044G06N 3/0442G06N 3/045G06N 3/084G06N 3/09G10L 21/0272G10L 25/30+5 more

AI classification

Speech1.00
Machine learning1.00
Natural language1.00
AI hardware1.00
Knowledge representation0.01
Vision0.00
Planning0.00
Evolutionary computation0.00

Ownership

MICROSOFT TECHNOLOGY LICENSING, LLC

assignment · 440450677

Assignors

CHEN, ZHUO, GONG, YIFAN, WANG, HUAMING, LI, JINYU, XIAO, XIONG, YOSHIOKA, TAKUYA, WANG, ZHENGHAO

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC