METHOD AND SYSTEM FOR DATA MINING OF VERY LARGE SPATIAL DATASETS USING VERTICAL SET INNER PRODUCTS
Patent №
US 7,836,090
Granted
2010-11-16
Filed 2007
Owner
NORTH DAKOTA STATE UNIVERSITY
+1 more
Lab
—
AI components
5
ml · vision · kr · planning · hardware
Assignment
Recorded
Dataset
AIPD
2023_r1 edition
Application
11791004
A system and method for performing and accelerating cluster analysis of large data sets is presented. The data set is formatted into binary bit Sequential (bSQ) format and then structured into a Peano Count tree (P-tree) format which represents a lossless tree representation of the original data. A P-tree algebra is defined and used to formulate a vertical set inner product (VSIP) technique that can be used to efficiently and scalably measure the mean value and total variation of a set about a fixed point in the large dataset. The set can be any projected subspace of any vector space, including oblique sub spaces. The VSIPs are used to determine the closeness of a point to a set of points in the large dataset making the VSIPs very useful in classification, clustering and outlier detection. One advantage is that the number of centroids (k) need not be pre-specified but are effectively determined. The high quality of the centroids makes them useful in partitioning clustering methods such as the k-means and the k-medoids clustering. The present invention also identifies the outliers.
AI classification
Ownership
NORTH DAKOTA STATE UNIVERSITY
assignment · 199000303
NDSU RESEARCH FOUNDATION
assignment · 200710363
Assignors
PERRIZO, WILLIAM K., ABIDIN, TAUFIK FUADI, PERERA, AMAL SHEHAN, SERAZI, MASUM
On an employer assignment, the assignors are typically the inventors.