EXTRACTING DATA CONTENT ITEMS USING TEMPLATE MATCHING

Patent №

US 7,765,236

Granted

2010-07-27

Filed 2007

Owner

MICROSOFT CORPORATION

AI components

3

nlp · kr · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11848987

Systems and methods for extracting data content items from a web page are provided. A template is created by labeling data content items of interest associated with a web page and generating a template Document Object Model (DOM) tree based on the labeled web page. DOM trees are also generated for additional web pages that contain data content items for which extraction may be desired. These DOM trees are compared to the template DOM tree to determine alignment there between. The aligned data content items may then be extracted from the additional web pages and indexed, as desired. Labeling the data content items of interest prior to generating a template DOM tree allows for the desired data content items to be specified and more accurately extracted from related and/or similarly structured web pages.

AI classification

Natural language1.00
AI hardware0.96
Knowledge representation0.70
Planning0.30
Vision0.01
Machine learning0.01
Evolutionary computation0.00
Speech0.00

Ownership

MICROSOFT CORPORATION

correct · 214120877

© 2026 NYSGPT2525 LLC