MULTILINGUAL SPEECH TRANSLATION WITH ADAPTIVE SPEECH SYNTHESIS AND ADAPTIVE PHYSIOGNOMY
Patent №
US 11,545,134
Granted
2023-01-03
Filed 2019
Owner
AMAZON TECHNOLOGIES, INC.
Lab
AI components
4
ml · nlp · speech · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16709792
Techniques for the generation of dubbed audio for an audio/video are described. An exemplary approach is to receive a request to generate dubbed speech for an audio/visual file; and in response to the request to: extract speech segments from an audio track of the audio/visual file associated with identified speakers; translate the extracted speech segments into a target language; determine a machine learning model per identified speaker, the trained machine learning models to be used to generate a spoken version of the translated, extracted speech segments based on the identified speaker; generate, per translated, extracted speech segment, a spoken version of the translated, extracted speech segments using a trained machine learning model that corresponds to the identified speaker of the translated, extracted speech segment and prosody information for the extracted speech segments; and replace the extracted speech segments from the audio track of the audio/visual file with the spoken versions spoken version of the translated, extracted speech segments to generate a modified audio track.
AI classification
Ownership
AMAZON TECHNOLOGIES, INC.
assignment · 513870001
Assignors
FEDERICO, MARCELLO, ENYEDI, ROBERT, AL-ONAIZAN, YASER, BARRA-CHICOTE, ROBERT, BREEN, ANDREW PAUL, GIRI, RITWIK, ISIK, MEHMET UMUT, KRISHNASWAMY, ARVINDH
On an employer assignment, the assignors are typically the inventors.