A Comparative Study of Resampling, Cost-Sensitive, and Ensemble Techniques for Handling Class Imbalance in Indonesian Financial Data
Handling class imbalance is a critical challenge in machine learning applications, particularly in financial domains where minority instances often represent significant anomalies such as fraud or audit risks. Various oversampling and undersampling methods were tested, alongside cost-sensitive adjustments and ensemble models including Random Forest, AdaBoost, Gradient Boosting, and XGBoost. The evaluation, based on 10-fold stratified cross-validation and performance metrics such as F1-score, ROC-AUC, and confusion matrix, highlights the superiority of a hybrid approach combining Borderline SMOTE and XGBoost. This configuration achieved near-perfect performance with F1-scores of 0.99 for both classes, demonstrating excellent discrimination and minimal error rates. The findings underscore the importance of method integration in imbalanced data scenarios and offer practical insights for model selection in real-world financial risk modeling.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex