METHODS, COMPUTING DEVICES, AND STORAGE MEDIA FOR GENERATING TRAINING CORPUS

Patent №

US 11,348,571

Granted

2022-05-31

Filed 2020

Owner

BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.

Lab

AI components

4

ml · nlp · speech · kr

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16810070

The present disclosure provides methods, computing devices, and storage media for generating a training corpus. The method includes: mining out pieces of data from user behavior logs associated with a target application, each piece of data including a first behavior log and a second behavior log, the first behavior log including a user speech and a corresponding speech recognition result, the second behavior log belonging to the same user as the first behavior log and time-dependent with the first behavior log; and determining the user speech and the corresponding speech recognition result in each piece of data as a positive feedback sample or a negative feedback sample, based on the first behavior log and the second behavior log.

Machine learningNatural languageSpeechKnowledge representationG10L 15/063G10L 15/26G10L 15/22G10L 25/63G10L 2015/0635G10L 2015/225

AI classification

Natural language1.00
Speech1.00
Machine learning1.00
Knowledge representation0.85
Planning0.39
AI hardware0.33
Evolutionary computation0.09
Vision0.00

Ownership

BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.

assignment · 520270686

Assignors

DING, SHIQIANG, HUANG, JIZHOU, JIANG, ZHONGWEI, MA, WENTAO

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC