SELF-SUPERVISED AI-ASSISTED SOUND EFFECT GENERATION FOR SILENT VIDEO USING MULTIMODAL CLUSTERING

Patent №

US 11,615,312

Granted

2023-03-28

Filed 2020

Owner

SONY INTERACTIVE ENTERTAINMENT INC.

Lab

AI components

7

ml · nlp · vision · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16848499

An automated method, system, and computer readable medium for generating sound effect recommendations for visual input by training machine learning models that learn audio-visual correlations from a reference image or video, a positive audio signal, and a negative audio signal. A machine learning algorithm is used with a reference visual input, a positive audio signal input or a negative audio signal input to train a multimodal clustering neural network to output representations for the visual input and audio input as well as correlation scores between the audio and visual representations. The trained multimodal clustering neural network is configured to learn representations in such a way that the visual representation and positive audio representation have higher correlation scores than the visual representation and a negative audio representation or an unrelated audio representation.

Machine learningNatural languageVisionSpeechKnowledge representationPlanningAI hardwareG06F 16/65G06N 3/084G06F 16/68G06N 3/044G06N 3/0442G06N 3/045G06N 3/0464G06N 3/088+4 more

AI classification

Machine learning1.00
Speech1.00
AI hardware1.00
Natural language0.99
Planning0.99
Vision0.99
Knowledge representation0.54
Evolutionary computation0.00

Ownership

SONY INTERACTIVE ENTERTAINMENT INC.

assignment · 523950070

Assignors

KRISHNAMURTHY, SUDHA

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC