Filter Methods for Feature Selection in Supervised Machine Learning Applications -- Review and Benchmark

The amount of data for machine learning (ML) applications is constantly\ngrowing. Not only the number of observations, especially the number of measured\nvariables (features) increases with ongoing digitization. Selecting the most\nappropriate features for predictive modeling is an important lever for the\nsuccess of ML applications in business and research. Feature selection methods\n(FSM) that are independent of a certain ML algorithm - so-called filter methods\n- have been numerously suggested, but little guidance for researchers and\nquantitative modelers exists to choose appropriate approaches for typical ML\nproblems. This review synthesizes the substantial literature on feature\nselection benchmarking and evaluates the performance of 58 methods in the\nwidely used R environment. For concrete guidance, we consider four typical\ndataset scenarios that are challenging for ML models (noisy, redundant,\nimbalanced data and cases with more features than observations). Drawing on the\nexperience of earlier benchmarks, which have considered much fewer FSMs, we\ncompare the performance of the methods according to four criteria (predictive\nperformance, number of relevant features selected, stability of the feature\nsets and runtime). We found methods relying on the random forest approach, the\ndouble input symmetrical relevance filter (DISR) and the joint impurity filter\n(JIM) were well-performing candidate methods for the given dataset scenarios.\n

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC