Knowledge graphs are freely aggregated, published, and edited in the Web of\ndata, and thus may overlap. Hence, a key task resides in aligning (or matching)\ntheir content. This task encompasses the identification, within an aggregated\nknowledge graph, of nodes that are equivalent, more specific, or weakly\nrelated. In this article, we propose to match nodes within a knowledge graph by\n(i) learning node embeddings with Graph Convolutional Networks such that\nsimilar nodes have low distances in the embedding space, and (ii) clustering\nnodes based on their embeddings, in order to suggest alignment relations\nbetween nodes of a same cluster. We conducted experiments with this approach on\nthe real world application of aligning knowledge in the field of\npharmacogenomics, which motivated our study. We particularly investigated the\ninterplay between domain knowledge and GCN models with the two following\nfocuses. First, we applied inference rules associated with domain knowledge,\nindependently or combined, before learning node embeddings, and we measured the\nimprovements in matching results. Second, while our GCN model is agnostic to\nthe exact alignment relations (e.g., equivalence, weak similarity), we observed\nthat distances in the embedding space are coherent with the ``strength'' of\nthese different relations (e.g., smaller distances for equivalences), letting\nus considering clustering and distances in the embedding space as a means to\nsuggest alignment relations in our case study.\n
Paper
References (38)
Scroll for more · 26 remaining