Efficient Twitter Sentiment Classification using Subjective Distant Supervision

As microblogging services like Twitter are becoming more and more influential\nin today's globalised world, its facets like sentiment analysis are being\nextensively studied. We are no longer constrained by our own opinion. Others\nopinions and sentiments play a huge role in shaping our perspective. In this\npaper, we build on previous works on Twitter sentiment analysis using Distant\nSupervision. The existing approach requires huge computation resource for\nanalysing large number of tweets. In this paper, we propose techniques to speed\nup the computation process for sentiment analysis. We use tweet subjectivity to\nselect the right training samples. We also introduce the concept of EFWS\n(Effective Word Score) of a tweet that is derived from polarity scores of\nfrequently used words, which is an additional heuristic that can be used to\nspeed up the sentiment classification with standard machine learning\nalgorithms. We performed our experiments using 1.6 million tweets. Experimental\nevaluations show that our proposed technique is more efficient and has higher\naccuracy compared to previously proposed methods. We achieve overall accuracies\nof around 80% (EFWS heuristic gives an accuracy around 85%) on a training\ndataset of 100K tweets, which is half the size of the dataset used for the\nbaseline model. The accuracy of our proposed model is 2-3% higher than the\nbaseline model, and the model effectively trains at twice the speed of the\nbaseline model.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC