Language-agnostic Multilingual Modeling Using Effective Script Normalization

Patent №

US 11,615,779

Granted

2023-03-28

Filed 2021

Owner

GOOGLE LLC

AI components

5

ml · nlp · speech · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

17152760

A method includes obtaining a plurality of training data sets each associated with a respective native language and includes a plurality of respective training data samples. For each respective training data sample of each training data set in the respective native language, the method includes transliterating the corresponding transcription in the respective native script into corresponding transliterated text representing the respective native language of the corresponding audio in a target script and associating the corresponding transliterated text in the target script with the corresponding audio in the respective native language to generate a respective normalized training data sample. The method also includes training, using the normalized training data samples, a multilingual end-to-end speech recognition model to predict speech recognition results in the target script for corresponding speech utterances spoken in any of the different native languages associated with the plurality of training data sets.

Machine learningNatural languageSpeechKnowledge representationAI hardwareG10L 15/005G06F 40/58G06N 3/044G06N 3/0442G06N 3/0455G06N 3/049G06N 3/084G06N 3/09+5 more

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
AI hardware1.00
Knowledge representation0.95
Vision0.24
Evolutionary computation0.00
Planning0.00

Ownership

GOOGLE LLC

assignment · 549780117

Assignors

DATTA, ARINDRIMA, RAMABHADRAN, BHUVANA, EMOND, JESSE, ROAK, BRIAN

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC