SYSTEM FOR AUTOMATICALLY MANAGING DUPLICATE DOCUMENTS WHEN CRAWLING DYNAMIC DOCUMENTS

Patent №

US 7,680,773

Granted

2010-03-16

Filed 2005

Owner

GOOGLE INC.

AI components

3

nlp · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11097687

A system of reducing the possibility of crawling duplicate document identifiers partitions a plurality of document identifiers into multiple clusters, each cluster having a cluster name and a set of document parameters. The system generates an equivalence rule for each cluster of document identifiers, the rule specifying which document parameters associated with the cluster are content-relevant. Next, the system groups each cluster of document identifiers into one or more equivalence classes in accordance with its associated equivalence rule, each equivalence class including one or more document identifiers that correspond to a document content and having a representative document identifier identifying the document content.

AI classification

Knowledge representation1.00
AI hardware0.98
Natural language0.96
Machine learning0.09
Planning0.01
Vision0.00
Evolutionary computation0.00
Speech0.00

Ownership

GOOGLE INC.

assignment · 160190024

Assignors

ACHARYA, ANURAG, JAIN, ARVIND, MUKHERJEE, ARUP

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC