DATA CLUSTERING USING ERROR-TOLERANT FREQUENT ITEM SETS

Patent №

US 6,567,936

Granted

2003-05-20

Filed 2000

Owner

MICROSOFT CORPORATION

AI components

3

ml · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

09500173

A generalization of frequent item sets to error-tolerant frequent item sets (ETF) is disclosed, together with its application in data clustering using error-tolerant frequent item sets to either build clusters or as an initialization technique for standard clustering algorithms. Efficient feasible computational algorithms for computing ETF's from very large databases is presented. In one embodiment, a method determines a plurality of weak ETF's, which are strongly tolerant of errors, and determines a plurality of strong ETF's therefrom, which are less tolerant of errors. The resulting clusters can be used as an initial model for a standard clustering approach, or may themselves be used as the end clusters. In one embodiment, the data covered by the strong clusters is removed from the data, and the process is repeated, until no more weak clusters can be found. Te invention includes methods for constructing ETF's from more general data types: data sets that include categorical discrete, continuous, and binary attributes.

Machine learningPlanningAI hardwareG06F 16/284G06F 2216/03Y10S 707/99932Y10S 707/99933Y10S 707/99934Y10S 707/99935Y10S 707/99936

AI classification

Machine learning1.00
Planning0.99
AI hardware0.99
Vision0.23
Natural language0.10
Knowledge representation0.04
Evolutionary computation0.00
Speech0.00

Ownership

MICROSOFT CORPORATION

assignment · 112050259

Assignors

YANG, CHENG, FAYYAD, USAMA M., BRADLEY, PAUL S.

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC