Patent US 5,680,481

Patent №

US 5,680,481

Granted

Owner

Lab

AI components

4

ml · nlp · vision · speech

Assignment

None on record

Dataset

AIPD

2023_r1 edition

Application

08488840

A facial feature extraction method and apparatus uses the variation in light intensity (gray-scale) of a frontal view of a speaker's face. The sequence of video images are sampled and quantized into a regular array of 150.times.150 pixels that naturally form a coordinate system of scan lines and pixel position along a scan line. Left and right eye areas and a mouth are located by thresholding the pixel gray-scale and finding the centroids of the three areas. The line segment joining the eye area centroids is bisected at right angle to form an axis of symmetry. A straight line through the centroid of the mouth area that is at right angle to the axis of symmetry constitutes the mouth line. Pixels along the mouth line and the axis of symmetry in the vicinity of the mouth area form a horizontal and vertical gray-scale profile, respectively. The profiles could be used as feature vectors but it is more efficient to select peaks and valleys (maximas and minimas) of the profile that correspond to the important physiological speech features such as lower and upper lip, mouth corner, and mouth area positions and pixel values and their time derivatives as visual vector components. Time derivatives are estimated by pixel position and value changes between video image frames. A speech recognition system uses the visual feature vector in combination with a concomitant acoustic vector as inputs to a time-delay neural network.

Machine learningNatural languageVisionSpeechG06V 40/171G06N 3/049G06V 30/248G06V 40/20G10L 15/25G10L 15/16

AI classification

Vision1.00
Speech1.00
Machine learning1.00
Natural language0.84
AI hardware0.01
Knowledge representation0.00
Evolutionary computation0.00
Planning0.00
© 2026 NYSGPT2525 LLC