NEAR-DUPLICATE DOCUMENT DETECTION FOR WEB CRAWLING

Patent №

US 8,548,972

Granted

2013-10-01

Filed 2012

Owner

Lab

AI components

5

ml · nlp · vision · kr · planning

Assignment

None on record

Dataset

AIPD

2023_r1 edition

Application

13422130

A system generates a hash value for a fetched document and compares the hash value with a set of stored hash values to identify ones of the stored hash values with a sequence of bit positions, less than all of the bit positions, that match a corresponding sequence of bit positions of the hash value. The system also determines whether any of the identified hash values are substantially similar to the hash value and identify the fetched document as a near-duplicate of another document when one of the identified hash values is substantially similar to the hash value.

AI classification

Planning1.00
Knowledge representation0.98
Natural language0.89
Vision0.60
Machine learning0.51
Speech0.05
AI hardware0.02
Evolutionary computation0.00
© 2026 NYSGPT2525 LLC