SYSTEM AND METHOD FOR IDENTIFYING USELESS DOCUMENTS

Patent №

US 6,397,211

Granted

2002-05-28

Filed 2000

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

5

ml · nlp · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

09476943

A system and method are disclosed for identifying useless or insignificant documents in a document hit list assembled from documents stored in one or more document collection databases. A search engine is used to compose the document hit list based on a query presented by a user. A text extraction algorithm run by a processor is then used to process the documents identified by the document hit list to produce a table of terms and their corresponding collection-level importance ranking called the IQ or Information Quotient. The text algorithm also produces a table of the most important terms per document. The documents are also scanned independently and a table of documents with filenames and lengths is also produced. A summarizing text algorithm is also run by a processor against the documents of the document hit list to produce a table of terms having a high tf*idf value for each document. All of the tables are stored in a relational database, which allows the system of the present invention to generate a table of terms per document ranked by decreasing IQ. To determine whether a document is useful or useless, the table of terms and IQs, the table of most important terms per document, the table of documents with filename and lengths, and the table of high tf*idf values are examined.

AI classification

Natural language1.00
AI hardware0.99
Knowledge representation0.99
Planning0.75
Machine learning0.72
Vision0.03
Speech0.00
Evolutionary computation0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 106140538

Assignors

COOPER, JAMES W.

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC