PHOTO-REALISTIC SYNTHESIS OF IMAGE SEQUENCES WITH LIP MOVEMENTS SYNCHRONIZED WITH SPEECH

Patent №

US 9,728,203

Granted

2017-08-08

Filed 2011

Owner

MICROSOFT CORPORATION

AI components

7

ml · nlp · vision · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

13098488

Audiovisual data of an individual reading a known script is obtained and stored in an audio library and an image library. The audiovisual data is processed to extract feature vectors used to train a statistical model. An input audio feature vector corresponding to desired speech with which a synthesized image sequence will be synchronized is provided. The statistical model is used to generate a trajectory of visual feature vectors that corresponds to the input audio feature vector. These visual feature vectors are used to identify a matching image sequence from the image library. The resulting sequence of images, concatenated from the image library, provides a photorealistic image sequence with lip movements synchronized with the desired speech.

AI classification

Vision1.00
Natural language1.00
Machine learning1.00
Knowledge representation1.00
AI hardware0.99
Speech0.94
Planning0.92
Evolutionary computation0.00

Ownership

MICROSOFT CORPORATION

assignment · 262060112

Assignors

WANG, LIJUAN, SOONG, FRANK

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC