Learning to Rank Anomalies: Scalar Performance Criteria and Maximization of Two-Sample Rank Statistics
The ability to collect and store ever more massive databases has been\naccompanied by the need to process them efficiently. In many cases, most\nobservations have the same behavior, while a probable small proportion of these\nobservations are abnormal. Detecting the latter, defined as outliers, is one of\nthe major challenges for machine learning applications (e.g. in fraud detection\nor in predictive maintenance). In this paper, we propose a methodology\naddressing the problem of outlier detection, by learning a data-driven scoring\nfunction defined on the feature space which reflects the degree of abnormality\nof the observations. This scoring function is learnt through a well-designed\nbinary classification problem whose empirical criterion takes the form of a\ntwo-sample linear rank statistics on which theoretical results are available.\nWe illustrate our methodology with preliminary encouraging numerical\nexperiments.\n