PROCESSING SPEECH SIGNALS OF A USER TO GENERATE A VISUAL REPRESENTATION OF THE USER
Patent №
US 11,568,864
Granted
2023-01-31
Filed 2019
Owner
CARNEGIE MELLON UNIVERSITY
AI components
5
ml · nlp · vision · speech · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16539701
A computing system for generating image data representing a speaker's face includes a detection device configured to route data representing a voice signal to one or more processors and a data processing device comprising the one or more processors configured to generate a representation of a speaker that generated the voice signal in response to receiving the voice signal. The data processing device executes a voice embedding function to generate a feature vector from the voice signal representing one or more signal features of the voice signal, maps a signal feature of the feature vector to a visual feature of the speaker by a modality transfer function specifying a relationship between the visual feature of the speaker and the signal feature of the feature vector; and generates a visual representation of at least a portion of the speaker based on the mapping, the visual representation comprising the visual feature.
AI classification
Ownership
CARNEGIE MELLON UNIVERSITY
assignment · 504170586
Assignors
SINGH, RITA
On an employer assignment, the assignors are typically the inventors.