END-TO-END AUTOMATIC SPEECH RECOGNITION SYSTEM FOR BOTH CONVERSATIONAL AND COMMAND-AND-CONTROL SPEECH
Patent №
US 12,223,953
Granted
2025-02-11
Filed 2022
Owner
NUANCE COMMUNICATIONS, INC.
Lab
—
AI components
0
Assignment
None on record
Dataset
AIPD
Application
17737587
A contextual end-to-end automatic speech recognition (ASR) system includes: an audio encoder configured to process input audio signal to produce as output encoded audio signal; a bias encoder configured to produce as output at least one bias entry corresponding to a word to bias for recognition by the ASR system; a transcription token probability prediction network configured to produce as output a probability of a selected transcription token, based at least in part on the output of the bias encoder and the output of the audio encoder; a first attention mechanism configured to receive the at least one bias entry and determine whether the at least one bias entry is suitable to be transcribed at a specific moment of an ongoing transcription; and a second attention mechanism configured to produce prefix penalties for restricting the first attention mechanism to only entries fitting a current transcription context.
Ownership
NUANCE COMMUNICATIONS, INC.