PROCESSING SPEECH SIGNALS OF A USER TO GENERATE A VISUAL REPRESENTATION OF THE USER

Patent №

US 11,568,864

Granted

2023-01-31

Filed 2019

Owner

CARNEGIE MELLON UNIVERSITY

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16539701

A computing system for generating image data representing a speaker's face includes a detection device configured to route data representing a voice signal to one or more processors and a data processing device comprising the one or more processors configured to generate a representation of a speaker that generated the voice signal in response to receiving the voice signal. The data processing device executes a voice embedding function to generate a feature vector from the voice signal representing one or more signal features of the voice signal, maps a signal feature of the feature vector to a visual feature of the speaker by a modality transfer function specifying a relationship between the visual feature of the speaker and the signal feature of the feature vector; and generates a visual representation of at least a portion of the speaker based on the mapping, the visual representation comprising the visual feature.

Machine learningNatural languageVisionSpeechAI hardwareG10L 15/22G10L 21/10G06T 11/60G10L 13/00G10L 15/02G10L 15/26G10L 2021/105

AI classification

Speech1.00
Natural language1.00
Vision1.00
AI hardware1.00
Machine learning0.96
Knowledge representation0.06
Planning0.01
Evolutionary computation0.00

Ownership

CARNEGIE MELLON UNIVERSITY

assignment · 504170586

Assignors

SINGH, RITA

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC