METHOD AND APPARATUS FOR EFFICIENT IDENTIFICATION OF DUPLICATE AND NEAR-DUPLICATE DOCUMENTS AND TEXT SPANS USING HIGH-DISCRIMINABILITY TEXT FRAGMENTS

Patent №

US 6,978,419

Granted

2005-12-20

Filed 2000

Owner

JUSTSYSTEM CORPORATION

Lab

AI components

3

nlp · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

09713733

Disclosed is a computer-assisted method for finding duplicate or near-duplicate documents or text spans within a document collection by using high-discriminability text fragments. Distinctive features of the documents or text spans are identified. For each pair of documents or text spans with at least one distinctive feature in common, the distinctive features of each document or text span are compared to determine whether the pair is duplicates or near-duplicates. An apparatus for performing this computer-assisted method is also disclosed.

AI classification

Natural language1.00
AI hardware0.94
Knowledge representation0.76
Evolutionary computation0.40
Machine learning0.22
Vision0.12
Planning0.00
Speech0.00

Ownership

JUSTSYSTEM CORPORATION

assignment · 112990352

Assignors

KANTROWITZ, MARK

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC