Sentiment analysis in Bengali via transfer learning using multi-lingual BERT

Sentiment analysis (SA) in Bengali is challenging due to this Indo-Aryan\nlanguage's highly inflected properties with more than 160 different inflected\nforms for verbs and 36 different forms for noun and 24 different forms for\npronouns. The lack of standard labeled datasets in the Bengali domain makes the\ntask of SA even harder. In this paper, we present manually tagged 2-class and\n3-class SA datasets in Bengali. We also demonstrate that the multi-lingual BERT\nmodel with relevant extensions can be trained via the approach of transfer\nlearning over those novel datasets to improve the state-of-the-art performance\nin sentiment classification tasks. This deep learning model achieves an\naccuracy of 71\\% for 2-class sentiment classification compared to the current\nstate-of-the-art accuracy of 68\\%. We also present the very first Bengali SA\nclassifier for the 3-class manually tagged dataset, and our proposed model\nachieves an accuracy of 60\\%. We further use this model to analyze the\nsentiment of public comments in the online daily newspaper. Our analysis shows\nthat people post negative comments for political or sports news more often,\nwhile the religious article comments represent positive sentiment. The dataset\nand code is publicly available at\nhttps://github.com/KhondokerIslam/Bengali\\_Sentiment.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC