Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional Regression
Random Forest (Breiman, 2001) is a successful and widely used regression and\nclassification algorithm. Part of its appeal and reason for its versatility is\nits (implicit) construction of a kernel-type weighting function on training\ndata, which can also be used for targets other than the original mean\nestimation. We propose a novel forest construction for multivariate responses\nbased on their joint conditional distribution, independent of the estimation\ntarget and the data model. It uses a new splitting criterion based on the MMD\ndistributional metric, which is suitable for detecting heterogeneity in\nmultivariate distributions. The induced weights define an estimate of the full\nconditional distribution, which in turn can be used for arbitrary and\npotentially complicated targets of interest. The method is very versatile and\nconvenient to use, as we illustrate on a wide range of examples. The code is\navailable as Python and R packages drf.\n
Paper
References (86)
Scroll for more · 38 remaining