BUILDING TRAINING DATA AND SIMILARITY RELATIONS FOR SEMANTIC SPACE

Patent №

US 10,755,183

Granted

2020-08-25

Filed 2017

Owner

EVERNOTE CORPORATION

Lab

AI components

4

nlp · vision · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

15416611

Selecting data from a source text corpus for training a semantic data analysis system includes selecting an item of the text corpus, validating the item, extracting at least one section of the item, determining a length of each of the at least one section of the item, and subdividing each of the sections having a length greater than a predetermined amount into a plurality of fragments that are deemed to be similar. The predetermined amount may be approximately twice a size of a fragment. A fragment may have approximately 100 words or between 40 and 60 words. Fragments from different items may be deemed to be dissimilar. Sections having a length less than the predetermined amount may be ignored. Validating the item may include parsing editorial notes and other accompanying data. The source text corpus may be Wikipedia. The item may be an article.

Natural languageVisionKnowledge representationAI hardwareG06F 40/131G06N 5/04G06F 16/334G06F 16/3344G06F 40/137G06F 40/166G06F 40/205G06F 40/216+4 more

AI classification

Natural language1.00
Vision1.00
AI hardware0.92
Knowledge representation0.57
Planning0.21
Machine learning0.16
Speech0.07
Evolutionary computation0.00

Ownership

EVERNOTE CORPORATION

assignment · 442820179

Assignors

LIVSHITZ, EUGENE, PASHINTSEV, ALEXANDER, GORBATOV, BORIS

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC