PREDICTING DATA FOR DOCUMENT ATTRIBUTES BASED ON AGGREGATED DATA FOR REPEATED URL PATTERNS
Patent №
US 8,645,367
Granted
2014-02-04
Filed 2010
Owner
GOOGLE INC.
AI components
2
nlp · kr
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
12719762
One or more hierarchies of string patterns are generated a plurality of URL strings according to a pattern extraction procedure. Repeated string patterns are selected from the generated hierarchies of string patterns. A URL class is defined for each of selected repeated string patterns. Each URL class is associated with a respective group of URL strings in the plurality of URL strings, where the respective group of URL strings contains a repeated string pattern that defines the URL class. Respective aggregated data is calculated for each URL class. The respective aggregated data is based on respective data of each respective document of each URL string in the group of URL strings associated with the URL class. Respective data for a respective document referenced by a lookup-URL is predicted based on respective aggregated data of one or more of the URL classes.
AI classification
Ownership
GOOGLE INC.
assignment · 247090040
Assignors
HAJAJ, NISSAN, ZHANG, CHI, WU, CHANGXUN, GROSS, ERIK
On an employer assignment, the assignors are typically the inventors.