Abstract We propose a semi-supervised approach towards anomaly detection in multivariate categorical data. Our goal is to learn a model that can distinguish the anomalous data, given a small set of training data from the normal class. To this end, our approach learns the probability distribution of normal instances with the assumption that the categorical data are generated from a continuous latent space. Gaussian process is adopted to construct the generative model. As a non-parametric Bayesian model, Gaussian process can adapt its model complexity according to the data size. Hence, our approach can be effective when the training dataset is small. Comprehensive experiments over different benchmarks clearly demonstrate the effectiveness of our approach.
Paper
Full text
Latent Gaussian process for anomaly detection in categorical data
Semantic Scholar · Computer Science · 2021
Abstract
Abstract We propose a semi-supervised approach towards anomaly detection in multivariate categorical data. Our goal is to learn a model that can distinguish the anomalous data, given a small set of training data from the normal class. To this end, our approach learns the probability distribution of normal instances with the assumption that the categorical data are generated from a continuous latent space. Gaussian process is adopted to construct the generative model. As a non-parametric Bayesian model, Gaussian process can adapt its model complexity according to the data size. Hence, our approach can be effective when the training dataset is small. Comprehensive experiments over different benchmarks clearly demonstrate the effectiveness of our approach.