Multilingual Culture-Independent Word Analogy Datasets

In text processing, deep neural networks mostly use word embeddings as an input. Embeddings have to ensure that relations between words are reflected through distances in a high-dimensional numeric space. To compare the quality of different text embeddings, typically, we use benchmark datasets. We present a collection of such datasets for the word analogy task in nine languages: Croatian, English, Estonian, Finnish, Latvian, Lithuanian, Russian, Slovenian, and Swedish. We designed the monolingual analogy task to be much more culturally independent and also constructed cross-lingual analogy datasets for the involved languages. We present basic statistics of the created datasets and their initial evaluation using fastText embeddings.

Paper

References (24)

08Here comes a link to the repository
09genitive to dative, a genitive noun case in relation to the dative noun case in respective languagesSlovene ceste : cesti: singular is used for all words, except ”human” (or equivalent in other languages), which appears in both singular and plural; in Finnish and Estonian, dative has been replaced with the alla-tive case; the category is not applicable to Swedish and English,
10family, a male family member in relation to an equivalent female member
11capitals and countriesParis : France
12FastText evaluation scores in % of correctly predicted relation pairs, i.e. how often was the vector d among the 10 closest vectors to the vector b − a + cTable

Scroll for more · 12 remaining

Similar papers

© 2026 NYSGPT2525 LLC