A Method for Handling Multi-class Imbalanced Data by Geometry based Information Sampling and Class Prioritized Synthetic Data Generation (GICaPS)
This paper looks into the problem of handling imbalanced data in a\nmulti-label classification problem. The problem is solved by proposing two\nnovel methods that primarily exploit the geometric relationship between the\nfeature vectors. The first one is an undersampling algorithm that uses angle\nbetween feature vectors to select more informative samples while rejecting the\nless informative ones. A suitable criterion is proposed to define the\ninformativeness of a given sample. The second one is an oversampling algorithm\nthat uses a generative algorithm to create new synthetic data that respects all\nclass boundaries. This is achieved by finding \\emph{no man's land} based on\nEuclidean distance between the feature vectors. The efficacy of the proposed\nmethods is analyzed by solving a generic multi-class recognition problem based\non mixture of Gaussians. The superiority of the proposed algorithms is\nestablished through comparison with other state-of-the-art methods, including\nSMOTE and ADASYN, over ten different publicly available datasets exhibiting\nhigh-to-extreme data imbalance. These two methods are combined into a single\ndata processing framework and is labeled as ``GICaPS'' to highlight the role of\ngeometry-based information (GI) sampling and Class-Prioritized Synthesis (CaPS)\nin dealing with multi-class data imbalance problem, thereby making a novel\ncontribution in this field.\n