SYSTEMS AND METHODS FOR NEURAL VOICE CLONING WITH A FEW SAMPLES

Patent №

US 11,238,843

Granted

2022-02-01

Filed 2018

Owner

BAIDU USA LLC

Lab

AI components

5

ml · nlp · vision · speech · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16143330

Voice cloning is a highly desired capability for personalized speech interfaces. Neural network-based speech synthesis has been shown to generate high quality speech for a large number of speakers. Neural voice cloning systems that take a few audio samples as input are presented herein. Two approaches, speaker adaptation and speaker encoding, are disclosed. Speaker adaptation embodiments are based on fine-tuning a multi-speaker generative model with a few cloning samples. Speaker encoding embodiments are based on training a separate model to directly infer a new speaker embedding from cloning audios, which is used in or with a multi-speaker generative model. Both approaches achieve good performance in terms of naturalness of the speech and its similarity to original speaker—even with very few cloning audios.

AI classification

Machine learning1.00
Speech1.00
Natural language1.00
AI hardware1.00
Vision1.00
Knowledge representation0.00
Planning0.00
Evolutionary computation0.00

Ownership

BAIDU USA LLC

assignment · 471400756

Assignors

ARIK, SERCAN O, CHEN, JITONG, PENG, KAINAN, PING, WEI, ZHOU, YANQI

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC