ANUBHUTI: A Comprehensive Corpus For Sentiment Analysis In Bangla Regional Languages

Sentiment analysis is a core task in Natural Language Processing (NLP) that focuses on identifying emotions expressed in text. Although widely studied in high-resource languages, Bangla, especially its regional dialects remains underexplored due to the lack of annotated resources. To address this gap, we introduce ANUBHUTI, a dialect-sensitive dataset containing 4,500 sentences equally distributed across three Bangla dialects: Noakhali, Sylhet, and Chittagong. The dataset was manually translated by native speakers and annotated with multiclass thematic labels (Political, Religious, Neutral) and multilabel emotions (Anger, Contempt, Disgust, Enjoyment, Fear, Sadness, Surprise). We evaluate classical baselines and transformer-based models (BanglaBERT, BanglaBERT Base, and mBERT) for multilabel emotion classification. Results show that transformer models, particularly BanglaBERT, consistently outperform traditional approaches across all dialects, highlighting the effectiveness of language-specific pretraining for dialectal sentiment analysis. ANUBHUTI provides a valuable benchmark for advancing sentiment analysis in Bangla regional dialects.

Paper

References (22)

Scroll for more · 10 remaining

Similar papers

© 2026 NYSGPT2525 LLC