A Simple and Efficient Ensemble Classifier Combining Multiple Neural Network Models on Social Media Datasets in Vietnamese
Text classification is a popular topic of natural language processing, which\nhas currently attracted numerous research efforts worldwide. The significant\nincrease of data in social media requires the vast attention of researchers to\nanalyze such data. There are various studies in this field in many languages\nbut limited to the Vietnamese language. Therefore, this study aims to classify\nVietnamese texts on social media from three different Vietnamese benchmark\ndatasets. Advanced deep learning models are used and optimized in this study,\nincluding CNN, LSTM, and their variants. We also implement the BERT, which has\nnever been applied to the datasets. Our experiments find a suitable model for\nclassification tasks on each specific dataset. To take advantage of single\nmodels, we propose an ensemble model, combining the highest-performance models.\nOur single models reach positive results on each dataset. Moreover, our\nensemble model achieves the best performance on all three datasets. We reach\n86.96% of F1- score for the HSD-VLSP dataset, 65.79% of F1-score for the\nUIT-VSMEC dataset, 92.79% and 89.70% for sentiments and topics on the UIT-VSFC\ndataset, respectively. Therefore, our models achieve better performances as\ncompared to previous studies on these datasets.\n