The identification of anomalous overdensities in data - group or collective\nanomaly detection - is a rich problem with a large number of real world\napplications. However, it has received relatively little attention in the\nbroader ML community, as compared to point anomalies or other types of single\ninstance outliers. One reason for this is the lack of powerful benchmark\ndatasets. In this paper, we first explain how, after the Nobel-prize winning\ndiscovery of the Higgs boson, unsupervised group anomaly detection has become a\nnew frontier of fundamental physics (where the motivation is to find new\nparticles and forces). Then we propose a realistic synthetic benchmark dataset\n(LHCO2020) for the development of group anomaly detection algorithms. Finally,\nwe compare several existing statistically-sound techniques for unsupervised\ngroup anomaly detection, and demonstrate their performance on the LHCO2020\ndataset.\n
Paper
References (43)
Scroll for more · 31 remaining