Method and System of Extracting Web Page Information

Patent №

US 9,767,211

Granted

2017-09-19

Filed 2015

Owner

ALIBABA GROUP HOLDING LIMITED

Lab

AI components

2

nlp · kr

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

14697696

A method of extracting web page information includes analyzing a document object model (DOM) structure of a sample page to obtain a position of information to be extracted. A node corresponding to the position of the information to be extracted is rendered in the DOM structure as a target node. Starting from the target node, relative position information is traversed recursively until the root node is found to create candidate paths. The candidate paths are rendered as a path set. A DOM structure of a page to be extracted is analyzed, information is located in the DOM structure of the page starting from the root node in the path set, and an extracted node candidate set is obtained. A node having highest robustness from the extracted node candidate set is selected to be a final extracted node and extracted information is obtained using the extracted node.

Natural languageKnowledge representationG06F 40/143G06F 16/80G06F 16/972G06F 40/103

AI classification

Natural language1.00
Knowledge representation1.00
Machine learning0.13
AI hardware0.02
Vision0.01
Evolutionary computation0.01
Planning0.00
Speech0.00

Ownership

ALIBABA GROUP HOLDING LIMITED

assignment · 355090889

Assignors

CAI, BOYANG, QIANG, QI

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC