A Rigorous Information-Theoretic Definition of Redundancy and Relevancy in Feature Selection Based on (Partial) Information Decomposition

Selecting a minimal feature set that is maximally informative about a target\nvariable is a central task in machine learning and statistics. Information\ntheory provides a powerful framework for formulating feature selection\nalgorithms -- yet, a rigorous, information-theoretic definition of feature\nrelevancy, which accounts for feature interactions such as redundant and\nsynergistic contributions, is still missing. We argue that this lack is\ninherent to classical information theory which does not provide measures to\ndecompose the information a set of variables provides about a target into\nunique, redundant, and synergistic contributions. Such a decomposition has been\nintroduced only recently by the partial information decomposition (PID)\nframework. Using PID, we clarify why feature selection is a conceptually\ndifficult problem when approached using information theory and provide a novel\ndefinition of feature relevancy and redundancy in PID terms. From this\ndefinition, we show that the conditional mutual information (CMI) maximizes\nrelevancy while minimizing redundancy and propose an iterative, CMI-based\nalgorithm for practical feature selection. We demonstrate the power of our\nCMI-based algorithm in comparison to the unconditional mutual information on\nbenchmark examples and provide corresponding PID estimates to highlight how PID\nallows to quantify information contribution of features and their interactions\nin feature-selection problems.\n

Paper

References (85)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC