EFFICIENT INDEXING OF DOCUMENTS WITH SIMILAR CONTENT

Patent №

US 8,554,561

Granted

2013-10-08

Filed 2012

Owner

Lab

AI components

4

nlp · kr · planning · evo

Assignment

None on record

Dataset

AIPD

2023_r1 edition

Application

13571316

A computer system comprising one or more processors and memory groups a set of documents into a plurality of clusters. Each cluster includes one or more documents of the set of documents and a respective cluster of documents of the plurality of clusters includes respective cluster data corresponding to a plurality of documents including a first document and a second document. The computer system determines that the second document includes duplicate data that is duplicative of corresponding data in the first document, identifies a respective subset of the respective cluster data that excludes at least a subset of the duplicate data, and generates an index of the respective subset of the respective cluster data.

AI classification

Natural language1.00
Planning0.95
Evolutionary computation0.56
Knowledge representation0.55
Machine learning0.47
AI hardware0.14
Vision0.00
Speech0.00
© 2026 NYSGPT2525 LLC