CSMOUTE: Combined Synthetic Oversampling and Undersampling Technique for Imbalanced Data Classification

In this paper we propose a novel data-level algorithm for handling data\nimbalance in the classification task, Synthetic Majority Undersampling\nTechnique (SMUTE). SMUTE leverages the concept of interpolation of nearby\ninstances, previously introduced in the oversampling setting in SMOTE.\nFurthermore, we combine both in the Combined Synthetic Oversampling and\nUndersampling Technique (CSMOUTE), which integrates SMOTE oversampling with\nSMUTE undersampling. The results of the conducted experimental study\ndemonstrate the usefulness of both the SMUTE and the CSMOUTE algorithms,\nespecially when combined with more complex classifiers, namely MLP and SVM, and\nwhen applied on datasets consisting of a large number of outliers. This leads\nus to a conclusion that the proposed approach shows promise for further\nextensions accommodating local data characteristics, a direction discussed in\nmore detail in the paper.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC