METHOD AND SYSTEM FOR ALIGNING NATURAL AND SYNTHETIC VIDEO TO SPEECH SYNTHESIS

Patent №

US 7,844,463

Granted

2010-11-30

Filed 2008

Owner

Lab

AI components

3

nlp · vision · hardware

Assignment

None on record

Dataset

AIPD

2023_r1 edition

Application

12193397

According to MPEG-4's TTS architecture, facial animation can be driven by two streams simultaneously—text and Facial Animation Parameters. A Text-To-Speech converter drives the mouth shapes of the face. An encoder sends Facial Animation Parameters to the face. The text input can include codes, or bookmarks, transmitted to the Text-to-Speech converter, which are placed between and inside words. The bookmarks carry an encoder time stamp. Due to the nature of text-to-speech conversion, the encoder time stamp does not relate to real-world time, and should be interpreted as a counter. The Facial Animation Parameter stream carries the same encoder time stamp found in the bookmark of the text. The system reads the bookmark and provides the encoder time stamp and a real-time time stamp. The facial animation system associates the correct facial animation parameter with the real-time time stamp using the encoder time stamp of the bookmark as a reference.

Natural languageVisionAI hardwareG06T 9/001G10L 13/00G10L 21/06H04N 21/2368H04N 21/4341

AI classification

Natural language1.00
AI hardware0.84
Vision0.63
Knowledge representation0.07
Machine learning0.00
Speech0.00
Evolutionary computation0.00
Planning0.00
© 2026 NYSGPT2525 LLC