Feature Selection Methods for Cost-Constrained Classification in Random Forests

Cost-sensitive feature selection describes a feature selection problem, where\nfeatures raise individual costs for inclusion in a model. These costs allow to\nincorporate disfavored aspects of features, e.g. failure rates of as measuring\ndevice, or patient harm, in the model selection process. Random Forests define\na particularly challenging problem for feature selection, as features are\ngenerally entangled in an ensemble of multiple trees, which makes a post hoc\nremoval of features infeasible. Feature selection methods therefore often\neither focus on simple pre-filtering methods, or require many Random Forest\nevaluations along their optimization path, which drastically increases the\ncomputational complexity. To solve both issues, we propose Shallow Tree\nSelection, a novel fast and multivariate feature selection method that selects\nfeatures from small tree structures. Additionally, we also adapt three standard\nfeature selection algorithms for cost-sensitive learning by introducing a\nhyperparameter-controlled benefit-cost ratio criterion (BCR) for each method.\nIn an extensive simulation study, we assess this criterion, and compare the\nproposed methods to multiple performance-based baseline alternatives on four\nartificial data settings and seven real-world data settings. We show that all\nmethods using a hyperparameterized BCR criterion outperform the baseline\nalternatives. In a direct comparison between the proposed methods, each method\nindicates strengths in certain settings, but no one-fits-all solution exists.\nOn a global average, we could identify preferable choices among our BCR based\nmethods. Nevertheless, we conclude that a practical analysis should never rely\non a single method only, but always compare different approaches to obtain the\nbest results.\n

Paper

References (36)

Scroll for more · 24 remaining

Similar papers

© 2026 NYSGPT2525 LLC