GENERATING TRAINING DATA FOR DISAMBIGUATION

Patent №

US 9,720,904

Granted

2017-08-01

Filed 2015

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

4

ml · nlp · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

14954636

A method for generating training data for disambiguation of an entity comprising a word or word string related to a topic to be analyzed includes acquiring sent messages by a user, each including at least one entity in a set of entities; organizing the messages and acquiring sets, each containing messages sent by each user; identifying a set of messages including different entities, greater than or equal to a first threshold value, and identifying a user corresponding to the identified set as a hot user; receiving an instruction indicating an object entity to be disambiguated; determining a likelihood of co-occurrence of each keyword and the object entity in sets of messages sent by hot users; and determining training data for the object entity on the basis of the likelihood of co-occurrence of each keyword and the object entity in the sets of messages sent by the hot users.

AI classification

Natural language1.00
Machine learning1.00
Knowledge representation1.00
AI hardware0.99
Planning0.43
Vision0.10
Evolutionary computation0.02
Speech0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 371700351

Assignors

IKAWA, YOHEI, SUZUKI, AKIKO

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC