Beyond the Selected Completely At Random Assumption for Learning from Positive and Unlabeled Data
Most positive and unlabeled data is subject to selection biases. The labeled\nexamples can, for example, be selected from the positive set because they are\neasier to obtain or more obviously positive. This paper investigates how\nlearning can be ena BHbled in this setting. We propose and theoretically\nanalyze an empirical-risk-based method for incorporating the labeling\nmechanism. Additionally, we investigate under which assumptions learning is\npossible when the labeling mechanism is not fully understood and propose a\npractical method to enable this. Our empirical analysis supports the\ntheoretical results and shows that taking into account the possibility of a\nselection bias, even when the labeling mechanism is unknown, improves the\ntrained classifiers.\n