Predicting and Explaining Behavioral Data with Structured Feature Space Decomposition

Modeling human behavioral data is challenging due to its scale, sparseness\n(few observations per individual), heterogeneity (differently behaving\nindividuals), and class imbalance (few observations of the outcome of\ninterest). An additional challenge is learning an interpretable model that not\nonly accurately predicts outcomes, but also identifies important factors\nassociated with a given behavior. To address these challenges, we describe a\nstatistical approach to modeling behavioral data called the structured\nsum-of-squares decomposition (S3D). The algorithm, which is inspired by\ndecision trees, selects important features that collectively explain the\nvariation of the outcome, quantifies correlations between the features, and\npartitions the subspace of important features into smaller, more homogeneous\nblocks that correspond to similarly-behaving subgroups within the population.\nThis partitioned subspace allows us to predict and analyze the behavior of the\noutcome variable both statistically and visually, giving a medium to examine\nthe effect of various features and to create explainable predictions. We apply\nS3D to learn models of online activity from large-scale data collected from\ndiverse sites, such as Stack Exchange, Khan Academy, Twitter, Duolingo, and\nDigg. We show that S3D creates parsimonious models that can predict outcomes in\nthe held-out data at levels comparable to state-of-the-art approaches, but in\naddition, produces interpretable models that provide insights into behaviors.\nThis is important for informing strategies aimed at changing behavior,\ndesigning social systems, but also for explaining predictions, a critical step\ntowards minimizing algorithmic bias.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC