Patent №
US 10,755,183
Granted
2020-08-25
Filed 2017
Owner
EVERNOTE CORPORATION
Lab
—
AI components
4
nlp · vision · kr · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
15416611
Selecting data from a source text corpus for training a semantic data analysis system includes selecting an item of the text corpus, validating the item, extracting at least one section of the item, determining a length of each of the at least one section of the item, and subdividing each of the sections having a length greater than a predetermined amount into a plurality of fragments that are deemed to be similar. The predetermined amount may be approximately twice a size of a fragment. A fragment may have approximately 100 words or between 40 and 60 words. Fragments from different items may be deemed to be dissimilar. Sections having a length less than the predetermined amount may be ignored. Validating the item may include parsing editorial notes and other accompanying data. The source text corpus may be Wikipedia. The item may be an article.
AI classification
Ownership
EVERNOTE CORPORATION
assignment · 442820179
Assignors
LIVSHITZ, EUGENE, PASHINTSEV, ALEXANDER, GORBATOV, BORIS
On an employer assignment, the assignors are typically the inventors.