GENERIC FRAMEWORK FOR LARGE-MARGIN MCE TRAINING IN SPEECH RECOGNITION

Patent №

US 8,423,364

Granted

2013-04-16

Filed 2007

Owner

MICROSOFT CORPORATION

AI components

5

ml · nlp · vision · speech · planning

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11708440

A method and apparatus for training an acoustic model are disclosed. A training corpus is accessed and converted into an initial acoustic model. Scores are calculated for a correct class and competitive classes, respectively, for each token given the initial acoustic model. Also, a sample-adaptive window bandwidth is calculated for each training token. From the calculated scores and the sample-adaptive window bandwidth values, loss values are calculated based on a loss function. The loss function, which may be derived from a Bayesian risk minimization viewpoint, can include a margin value that moves a decision boundary such that token-to-boundary distances for correct tokens that are near the decision boundary are maximized. The margin can either be a fixed margin or can vary monotonically as a function of algorithm iterations. The acoustic model is updated based on the calculated loss values. This process can be repeated until an empirical convergence is met.

AI classification

Machine learning1.00
Speech1.00
Vision1.00
Natural language0.97
Planning0.68
AI hardware0.43
Knowledge representation0.00
Evolutionary computation0.00

Ownership

MICROSOFT CORPORATION

assignment · 190400122

Assignors

YU, DONG, ACERO, ALEJANDRO, DENG, LI, HE, XIAODONG

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC