END-TO-END AUTOMATIC SPEECH RECOGNITION SYSTEM FOR BOTH CONVERSATIONAL AND COMMAND-AND-CONTROL SPEECH

Patent №

US 12,223,953

Granted

2025-02-11

Filed 2022

Owner

NUANCE COMMUNICATIONS, INC.

Lab

AI components

0

Assignment

None on record

Dataset

AIPD

Application

17737587

A contextual end-to-end automatic speech recognition (ASR) system includes: an audio encoder configured to process input audio signal to produce as output encoded audio signal; a bias encoder configured to produce as output at least one bias entry corresponding to a word to bias for recognition by the ASR system; a transcription token probability prediction network configured to produce as output a probability of a selected transcription token, based at least in part on the output of the bias encoder and the output of the audio encoder; a first attention mechanism configured to receive the at least one bias entry and determine whether the at least one bias entry is suitable to be transcribed at a specific moment of an ongoing transcription; and a second attention mechanism configured to produce prefix penalties for restricting the first attention mechanism to only entries fitting a current transcription context.

G10L 15/16G10L 15/197G06N 3/0455G06N 3/0442G06N 3/088G10L 15/30

Ownership

NUANCE COMMUNICATIONS, INC.

© 2026 NYSGPT2525 LLC