Comparison of Turkish Word Representations Trained on Different Morphological Forms

Increased popularity of different text representations has also brought many\nimprovements in Natural Language Processing (NLP) tasks. Without need of\nsupervised data, embeddings trained on large corpora provide us meaningful\nrelations to be used on different NLP tasks. Even though training these vectors\nis relatively easy with recent methods, information gained from the data\nheavily depends on the structure of the corpus language. Since the popularly\nresearched languages have a similar morphological structure, problems occurring\nfor morphologically rich languages are mainly disregarded in studies. For\nmorphologically rich languages, context-free word vectors ignore morphological\nstructure of languages. In this study, we prepared texts in morphologically\ndifferent forms in a morphologically rich language, Turkish, and compared the\nresults on different intrinsic and extrinsic tasks. To see the effect of\nmorphological structure, we trained word2vec model on texts which lemma and\nsuffixes are treated differently. We also trained subword model fastText and\ncompared the embeddings on word analogy, text classification, sentimental\nanalysis, and language model tasks.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC