SELF-SUPERVISED AI-ASSISTED SOUND EFFECT GENERATION FOR SILENT VIDEO USING MULTIMODAL CLUSTERING
Patent №
US 11,615,312
Granted
2023-03-28
Filed 2020
Owner
SONY INTERACTIVE ENTERTAINMENT INC.
Lab
—
AI components
7
ml · nlp · vision · speech · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
16848499
An automated method, system, and computer readable medium for generating sound effect recommendations for visual input by training machine learning models that learn audio-visual correlations from a reference image or video, a positive audio signal, and a negative audio signal. A machine learning algorithm is used with a reference visual input, a positive audio signal input or a negative audio signal input to train a multimodal clustering neural network to output representations for the visual input and audio input as well as correlation scores between the audio and visual representations. The trained multimodal clustering neural network is configured to learn representations in such a way that the visual representation and positive audio representation have higher correlation scores than the visual representation and a negative audio representation or an unrelated audio representation.
AI classification
Ownership
SONY INTERACTIVE ENTERTAINMENT INC.
assignment · 523950070
Assignors
KRISHNAMURTHY, SUDHA
On an employer assignment, the assignors are typically the inventors.