Unsupervised Anomaly Detection in Knowledge Graphs

Anomalies such as redundant, inconsistent, contradictory, and deficient values in a knowledge graph are unavoidable, as such graphs are often curated manually, or extracted using machine learning and natural language processing techniques. Therefore, anomaly detection in knowledge graphs is an essential task that contributes towards its quality. Although there are approaches to detect anomalies in knowledge graphs, they are either domain dependent, not scalable to large graphs, or they require substantial human intervention. In this preliminary research paper we propose a novel unsupervised feature-based approach to anomaly detection in knowledge graphs. We first characterize triples in a directed edge-labelled knowledge graph using a set of binary features, and then use a one-class Support Vector Machine (SVM) to classify these triples as normal or abnormal. After selecting the features that have the highest consistency with the SVM outcomes, we provide a visualization of the identified anomalies, and the list of anomalous triples, thus supporting non-technical domain experts to understand the anomalies present in a knowledge graph. We evaluate our approach on the four knowledge graphs YAGO-1, KBpedia, Wikidata, and DSKG. This evaluation demonstrates that our approach is well suited to identify anomalies in knowledge graphs in an unsupervised manner, independent from the domain of the knowledge graph being evaluated.

Paper

Full text

PDF

Unsupervised Anomaly Detection in Knowledge Graphs

Semantic Scholar · Computer Science · 2021

Abstract

Anomalies such as redundant, inconsistent, contradictory, and deficient values in a knowledge graph are unavoidable, as such graphs are often curated manually, or extracted using machine learning and natural language processing techniques. Therefore, anomaly detection in knowledge graphs is an essential task that contributes towards its quality. Although there are approaches to detect anomalies in knowledge graphs, they are either domain dependent, not scalable to large graphs, or they require substantial human intervention. In this preliminary research paper we propose a novel unsupervised feature-based approach to anomaly detection in knowledge graphs. We first characterize triples in a directed edge-labelled knowledge graph using a set of binary features, and then use a one-class Support Vector Machine (SVM) to classify these triples as normal or abnormal. After selecting the features that have the highest consistency with the SVM outcomes, we provide a visualization of the identified anomalies, and the list of anomalous triples, thus supporting non-technical domain experts to understand the anomalies present in a knowledge graph. We evaluate our approach on the four knowledge graphs YAGO-1, KBpedia, Wikidata, and DSKG. This evaluation demonstrates that our approach is well suited to identify anomalies in knowledge graphs in an unsupervised manner, independent from the domain of the knowledge graph being evaluated.

References (26)

Scroll for more · 14 remaining

Similar papers

© 2026 NYSGPT2525 LLC