Clustering Word Embeddings with Self-Organizing Maps. Application on LaRoSeDa -- A Large Romanian Sentiment Data Set
Romanian is one of the understudied languages in computational linguistics,\nwith few resources available for the development of natural language processing\ntools. In this paper, we introduce LaRoSeDa, a Large Romanian Sentiment Data\nSet, which is composed of 15,000 positive and negative reviews collected from\none of the largest Romanian e-commerce platforms. We employ two sentiment\nclassification methods as baselines for our new data set, one based on\nlow-level features (character n-grams) and one based on high-level features\n(bag-of-word-embeddings generated by clustering word embeddings with k-means).\nAs an additional contribution, we replace the k-means clustering algorithm with\nself-organizing maps (SOMs), obtaining better results because the generated\nclusters of word embeddings are closer to the Zipf's law distribution, which is\nknown to govern natural language. We also demonstrate the generalization\ncapacity of using SOMs for the clustering of word embeddings on another\nrecently-introduced Romanian data set, for text categorization by topic.\n
Paper
References (57)
Scroll for more · 38 remaining