METHOD AND SYSTEM FOR DATA MINING OF VERY LARGE SPATIAL DATASETS USING VERTICAL SET INNER PRODUCTS

Patent №

US 7,836,090

Granted

2010-11-16

Filed 2007

Owner

NORTH DAKOTA STATE UNIVERSITY

+1 more

Lab

AI components

5

ml · vision · kr · planning · hardware

Assignment

Recorded

Dataset

AIPD

2023_r1 edition

Application

11791004

A system and method for performing and accelerating cluster analysis of large data sets is presented. The data set is formatted into binary bit Sequential (bSQ) format and then structured into a Peano Count tree (P-tree) format which represents a lossless tree representation of the original data. A P-tree algebra is defined and used to formulate a vertical set inner product (VSIP) technique that can be used to efficiently and scalably measure the mean value and total variation of a set about a fixed point in the large dataset. The set can be any projected subspace of any vector space, including oblique sub spaces. The VSIPs are used to determine the closeness of a point to a set of points in the large dataset making the VSIPs very useful in classification, clustering and outlier detection. One advantage is that the number of centroids (k) need not be pre-specified but are effectively determined. The high quality of the centroids makes them useful in partitioning clustering methods such as the k-means and the k-medoids clustering. The present invention also identifies the outliers.

AI classification

Machine learning1.00
Vision1.00
Knowledge representation1.00
Planning0.99
AI hardware0.88
Natural language0.07
Speech0.00
Evolutionary computation0.00

Ownership

NORTH DAKOTA STATE UNIVERSITY

assignment · 199000303

NDSU RESEARCH FOUNDATION

assignment · 200710363

Assignors

PERRIZO, WILLIAM K., ABIDIN, TAUFIK FUADI, PERERA, AMAL SHEHAN, SERAZI, MASUM

On an employer assignment, the assignors are typically the inventors.

© 2026 NYSGPT2525 LLC