MULTILINGUAL SPEECH TRANSLATION WITH ADAPTIVE SPEECH SYNTHESIS AND ADAPTIVE PHYSIOGNOMY

Patent №

US 11,545,134

Granted

2023-01-03

Filed 2019

Owner

AMAZON TECHNOLOGIES, INC.

AI components

4

ml · nlp · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16709792

Techniques for the generation of dubbed audio for an audio/video are described. An exemplary approach is to receive a request to generate dubbed speech for an audio/visual file; and in response to the request to: extract speech segments from an audio track of the audio/visual file associated with identified speakers; translate the extracted speech segments into a target language; determine a machine learning model per identified speaker, the trained machine learning models to be used to generate a spoken version of the translated, extracted speech segments based on the identified speaker; generate, per translated, extracted speech segment, a spoken version of the translated, extracted speech segments using a trained machine learning model that corresponds to the identified speaker of the translated, extracted speech segment and prosody information for the extracted speech segments; and replace the extracted speech segments from the audio track of the audio/visual file with the spoken versions spoken version of the translated, extracted speech segments to generate a modified audio track.

Machine learningNatural languageSpeechAI hardwareG06F 40/44G10L 13/10G06F 18/214G06F 40/253G06F 40/47G06F 40/58G06V 40/161G06V 40/168+8 more

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
AI hardware0.94
Knowledge representation0.10
Evolutionary computation0.08
Vision0.02
Planning0.00

Ownership

AMAZON TECHNOLOGIES, INC.

assignment · 513870001

Assignors

FEDERICO, MARCELLO, ENYEDI, ROBERT, AL-ONAIZAN, YASER, BARRA-CHICOTE, ROBERT, BREEN, ANDREW PAUL, GIRI, RITWIK, ISIK, MEHMET UMUT, KRISHNASWAMY, ARVINDH

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC