LU-BZU at SemEval-2021 Task 2: Word2Vec and Lemma2Vec performance in Arabic Word-in-Context disambiguation

This paper presents a set of experiments to evaluate and compare between the\nperformance of using CBOW Word2Vec and Lemma2Vec models for Arabic\nWord-in-Context (WiC) disambiguation without using sense inventories or sense\nembeddings. As part of the SemEval-2021 Shared Task 2 on WiC disambiguation, we\nused the dev.ar-ar dataset (2k sentence pairs) to decide whether two words in a\ngiven sentence pair carry the same meaning. We used two Word2Vec models:\nWiki-CBOW, a pre-trained model on Arabic Wikipedia, and another model we\ntrained on large Arabic corpora of about 3 billion tokens. Two Lemma2Vec models\nwas also constructed based on the two Word2Vec models. Each of the four models\nwas then used in the WiC disambiguation task, and then evaluated on the\nSemEval-2021 test.ar-ar dataset. At the end, we reported the performance of\ndifferent models and compared between using lemma-based and word-based models.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC