A Comparison of Semantic Similarity Methods for Maximum Human Interpretability

The inclusion of semantic information in any similarity measures improves the\nefficiency of the similarity measure and provides human interpretable results\nfor further analysis. The similarity calculation method that focuses on\nfeatures related to the text's words only, will give less accurate results.\nThis paper presents three different methods that not only focus on the text's\nwords but also incorporates semantic information of texts in their feature\nvector and computes semantic similarities. These methods are based on\ncorpus-based and knowledge-based methods, which are: cosine similarity using\ntf-idf vectors, cosine similarity using word embedding and soft cosine\nsimilarity using word embedding. Among these three, cosine similarity using\ntf-idf vectors performed best in finding similarities between short news texts.\nThe similar texts given by the method are easy to interpret and can be used\ndirectly in other information retrieval applications.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC