TECHNIQUES FOR CORRECTING LINGUISTIC TRAINING BIAS IN TRAINING DATA

Patent №

US 11,373,090

Granted

2022-06-28

Filed 2018

Owner

TATA CONSULTANCY SERVICES LIMITED

Lab

AI components

7

ml · nlp · vision · speech · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

16134360

In automated assistant systems, a deep-learning model in form of a long short-term memory (LSTM) classifier is used for mapping questions to classes, with each class having a manually curated answer. A team of experts manually create the training data used to train this classifier. Relying on human curation often results in such linguistic training biases creeping into training data, since every individual has a specific style of writing natural language and uses some words in specific context only. Deep models end up learning these biases, instead of the core concept words of the target classes. In order to correct these biases, meaningful sentences are automatically generated using a generative model, and then used for training a classification model. For example, a variational autoencoder (VAE) is used as the generative model for generating novel sentences and a language model (LM) is utilized for selecting sentences based on likelihood.

Machine learningNatural languageVisionSpeechKnowledge representationPlanningAI hardwareG06F 16/3329G06N 3/08G06F 16/2455G06N 3/044G06N 3/0442G06N 3/045G06N 3/0455G06N 3/047+3 more

AI classification

Natural language1.00
Machine learning1.00
Vision1.00
Speech1.00
Knowledge representation1.00
AI hardware1.00
Planning0.96
Evolutionary computation0.00

Ownership

TATA CONSULTANCY SERVICES LIMITED

assignment · 469000921

Assignors

AGARWAL, PUNEET, PATIDAR, MAYUR, VIG, LOVEKESH, SHROFF, GAUTAM

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC