SINGING VOICE SEPARATION WITH DEEP U-NET CONVOLUTIONAL NETWORKS

Patent №

US 10,923,142

Granted

2021-02-16

Filed 2019

Owner

SPOTIFY AB

Lab

AI components

4

ml · nlp · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16242525

A system, method and computer product for training a neural network system. The method comprises applying an audio signal to the neural network system, the audio signal including a vocal component and a non-vocal component. The method also comprises comparing an output of the neural network system to a target signal, and adjusting at least one parameter of the neural network system to reduce a result of the comparing, for training the neural network system to estimate one of the vocal component and the non-vocal component. In one example embodiment, the system comprises a U-Net architecture. After training, the system can estimate vocal or instrumental components of an audio signal, depending on which type of component the system is trained to estimate.

Machine learningNatural languageSpeechAI hardwareG10H 1/0008G10L 25/81G06N 3/045G06N 3/0455G06N 3/0464G06N 3/08G06N 3/09G06N 5/046+9 more

AI classification

Speech1.00
Natural language1.00
Machine learning1.00
AI hardware0.99
Vision0.10
Planning0.01
Knowledge representation0.01
Evolutionary computation0.00

Ownership

SPOTIFY AB

assignment · 540240149

Assignors

JANSSON, ANDREAS SIMON THORE, SACKFIELD, ANGUS WILLIAM, SUNG, CHING

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC