Improving Sentiment Analysis over non-English Tweets using Multilingual Transformers and Automatic Translation for Data-Augmentation
Tweets are specific text data when compared to general text. Although\nsentiment analysis over tweets has become very popular in the last decade for\nEnglish, it is still difficult to find huge annotated corpora for non-English\nlanguages. The recent rise of the transformer models in Natural Language\nProcessing allows to achieve unparalleled performances in many tasks, but these\nmodels need a consequent quantity of text to adapt to the tweet domain. We\npropose the use of a multilingual transformer model, that we pre-train over\nEnglish tweets and apply data-augmentation using automatic translation to adapt\nthe model to non-English languages. Our experiments in French, Spanish, German\nand Italian suggest that the proposed technique is an efficient way to improve\nthe results of the transformers over small corpora of tweets in a non-English\nlanguage.\n
Paper
References (27)
Scroll for more · 15 remaining