Exploiting Vietnamese Social Media Characteristics for Textual Emotion Recognition in Vietnamese
Textual emotion recognition has been a promising research topic in recent\nyears. Many researchers aim to build more accurate and robust emotion detection\nsystems. In this paper, we conduct several experiments to indicate how data\npre-processing affects a machine learning method on textual emotion\nrecognition. These experiments are performed on the Vietnamese Social Media\nEmotion Corpus (UIT-VSMEC) as the benchmark dataset. We explore Vietnamese\nsocial media characteristics to propose different pre-processing techniques,\nand key-clause extraction with emotional context to improve the machine\nperformance on UIT-VSMEC. Our experimental evaluation shows that with\nappropriate pre-processing techniques based on Vietnamese social media\ncharacteristics, Multinomial Logistic Regression (MLR) achieves the best\nF1-score of 64.40%, a significant improvement of 4.66% over the CNN model built\nby the authors of UIT-VSMEC (59.74%).\n