TEXT-TO-SPEECH FROM MEDIA CONTENT ITEM SNIPPETS

Patent №

US 11,114,085

Granted

2021-09-07

Filed 2018

Owner

SPOTIFY AB

+1 more

Lab

AI components

2

nlp · speech

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16235776

A text-to-speech engine creates audio output that includes synthesized speech and one or more media content item snippets. The input text is obtained and partitioned into text sets. A track having lyrics that match a part of one of the text sets is identified. The location of the track's audio that contains the lyric is extracted based on forced alignment data. The extracted audio is combined with synthesized speech corresponding to the remainder of the input text to form audio output.

Natural languageSpeechG10L 13/00G10L 13/06G06F 16/685G10L 13/02G10L 13/04

AI classification

Natural language1.00
Speech1.00
Machine learning0.49
AI hardware0.16
Planning0.10
Knowledge representation0.03
Evolutionary computation0.00
Vision0.00

Ownership

SPOTIFY AB

assignment · 548280573

SPOTIFY USA INC.

correct · 626620199

© 2026 NYSGPT2525 LLC