TECHNIQUES FOR DETECTING DUPLICATE WEB PAGES

Patent №

US 7,698,317

Granted

2010-04-13

Filed 2007

Owner

YAHOO! INC.

Lab

AI components

6

ml · vision · kr · planning · evo · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11788505

Techniques are disclosed for detecting web pages with duplicate content. In one embodiment, a set of shingles is computed for each page of a group of pages. An aggregate set of shingles is determined based on the sets of shingles computed for the group of pages. A first subset from the aggregate set of shingles is determined by selecting, from the aggregate set, shingles whose frequencies in the aggregate set exceed a specified threshold. A modified set of shingles is generated for each page of the group of pages by removing, from the set of shingles for that page, any shingle included in the first subset. One or more duplicate pages in the group of pages are determined based at least in part on the modified sets of shingles generated for the group of pages.

AI classification

Knowledge representation1.00
Evolutionary computation0.99
AI hardware0.99
Planning0.90
Machine learning0.68
Vision0.59
Natural language0.17
Speech0.00

Ownership

YAHOO! INC.

assignment · 192790001

Assignors

SASTURKAR, AMIT, AHUJA, RAJAT, RAVIKUMAR, SHANMUGASUNDARAM, OFITSEROV, VLADIMIR

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC