PHOTO-REALISTIC SYNTHESIS OF IMAGE SEQUENCES WITH LIP MOVEMENTS SYNCHRONIZED WITH SPEECH
Patent №
US 9,728,203
Granted
2017-08-08
Filed 2011
Owner
MICROSOFT CORPORATION
Lab
AI components
7
ml · nlp · vision · speech · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
13098488
Audiovisual data of an individual reading a known script is obtained and stored in an audio library and an image library. The audiovisual data is processed to extract feature vectors used to train a statistical model. An input audio feature vector corresponding to desired speech with which a synthesized image sequence will be synchronized is provided. The statistical model is used to generate a trajectory of visual feature vectors that corresponds to the input audio feature vector. These visual feature vectors are used to identify a matching image sequence from the image library. The resulting sequence of images, concatenated from the image library, provides a photorealistic image sequence with lip movements synchronized with the desired speech.
AI classification
Ownership
MICROSOFT CORPORATION
assignment · 262060112
Assignors
WANG, LIJUAN, SOONG, FRANK
On an employer assignment, the assignors are typically the inventors.