Pairwise Multi-Class Document Classification for Semantic Relations between Wikipedia Articles

Many digital libraries recommend literature to their users considering the\nsimilarity between a query document and their repository. However, they often\nfail to distinguish what is the relationship that makes two documents alike. In\nthis paper, we model the problem of finding the relationship between two\ndocuments as a pairwise document classification task. To find the semantic\nrelation between documents, we apply a series of techniques, such as GloVe,\nParagraph-Vectors, BERT, and XLNet under different configurations (e.g.,\nsequence length, vector concatenation scheme), including a Siamese architecture\nfor the Transformer-based systems. We perform our experiments on a newly\nproposed dataset of 32,168 Wikipedia article pairs and Wikidata properties that\ndefine the semantic document relations. Our results show vanilla BERT as the\nbest performing system with an F1-score of 0.93, which we manually examine to\nbetter understand its applicability to other domains. Our findings suggest that\nclassifying semantic relations between documents is a solvable task and\nmotivates the development of recommender systems based on the evaluated\ntechniques. The discussions in this paper serve as first steps in the\nexploration of documents through SPARQL-like queries such that one could find\ndocuments that are similar in one aspect but dissimilar in another.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC