SEMI-SUPERVISED SYSTEM FOR MULTICHANNEL SOURCE ENHANCEMENT THROUGH CONFIGURABLE ADAPTIVE TRANSFORMATIONS AND DEEP NEURAL NETWORK
Patent №
US 10,347,271
Granted
2019-07-09
Filed 2016
Owner
CONEXANT SYSTEMS, LLC
Lab
—
AI components
3
ml · vision · speech
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
15368452
Various techniques are provided to perform enhanced automatic speech recognition. For example, a subband analysis may be performed that transforms time-domain signals of multiple audio channels in subband signals. An adaptive configurable transformation may also be performed to produce single or multichannel-based features whose values are correlated to an Ideal Binary Mask (IBM). An unsupervised Gaussian Mixture Model (GMM) model fitting the distribution of the features and producing posterior probabilities may also be performed, and the posteriors may be combined to produce deep neural network (DNN) feature vectors. A DNN may be provided that predicts oracle spectral gains from the input feature vectors. Spectral processing may be performed to produce an estimate of the target source time-frequency magnitudes from the mixtures and the output of the DNN. Subband synthesis may be performed to transform signals back to time-domain.
AI classification
Ownership
CONEXANT SYSTEMS, LLC
assignment · 430030102
Assignors
NESTA, FRANCESCO, ZHAO, XIANGYUAN, THORMUNDSSON, TRAUSTI
On an employer assignment, the assignors are typically the inventors.