SELF-ADAPTIVE WEB CRAWLING AND TEXT EXTRACTION

Patent №

US 10,922,366

Granted

2021-02-16

Filed 2018

Owner

INTERNATIONAL BUSINESS MACHINES CORPORATION

AI components

5

ml · nlp · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

15936666

A method, computer system, and a computer program product for crawling and extracting main content from a web page is provided. The present invention may include retrieving a HTML document associated with a web page. The present invention may then include identifying at least one entry point located in the retrieved HTML document by utilizing a self-adaptive entry point locator. The present invention may also include extracting a main content article associated with the retrieved HTML document based on the identified at least one entry point. The present invention may further include presenting the extracted main content associated with the retrieved HTML document to the user.

Machine learningNatural languageKnowledge representationPlanningAI hardwareG06F 16/951G06F 16/9535G06F 16/334G06F 16/986G06F 40/103G06F 40/279H04L 67/02

AI classification

Natural language1.00
Knowledge representation1.00
Machine learning1.00
Planning0.92
AI hardware0.85
Vision0.06
Evolutionary computation0.00
Speech0.00

Ownership

INTERNATIONAL BUSINESS MACHINES CORPORATION

assignment · 453960132

Assignors

HUANG, CHEN-YU, LEE, SHENG-WEI, LIN, JUNE-RAY, WU, CI-HAO, YANG, HSIEH-LUNG, YU, YING-CHEN

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC