OPEN INFORMATION EXTRACTION FROM THE WEB

Patent №

US 7,877,343

Granted

2011-01-25

Filed 2007

Owner

UNIVERSITY OF WASHINGTON

Lab

AI components

7

ml · nlp · vision · kr · planning · evo · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11695506

To implement open information extraction, a new extraction paradigm has been developed in which a system makes a single data-driven pass over a corpus of text, extracting a large set of relational tuples without requiring any human input. Using training data, a Self-Supervised Learner employs a parser and heuristics to determine criteria that will be used by an extraction classifier (or other ranking model) for evaluating the trustworthiness of candidate tuples that have been extracted from the corpus of text, by applying heuristics to the corpus of text. The classifier retains tuples with a sufficiently high probability of being trustworthy. A redundancy-based assessor assigns a probability to each retained tuple to indicate a likelihood that the retained tuple is an actual instance of a relationship between a plurality of objects comprising the retained tuple. The retained tuples comprise an extraction graph that can be queried for information.

AI classification

Natural language1.00
Machine learning1.00
Knowledge representation1.00
Planning1.00
Vision0.99
AI hardware0.99
Evolutionary computation0.91
Speech0.00

Ownership

UNIVERSITY OF WASHINGTON

assignment · 191210103

Assignors

CAFARELLA, MICHAEL J., BANKO, MICHELE, ETZIONI, OREN

On an employer assignment, the assignors are typically the inventors.

From the same owner

© 2026 NYSGPT2525 LLC