The recent surge of complex attention-based deep learning architectures has\nled to extraordinary results in various downstream NLP tasks in the English\nlanguage. However, such research for resource-constrained and morphologically\nrich Indian vernacular languages has been relatively limited. This paper\nproffers team SPPU\\_AKAH's solution for the TechDOfication 2020 subtask-1f:\nwhich focuses on the coarse-grained technical domain identification of short\ntext documents in Marathi, a Devanagari script-based Indian language. Availing\nthe large dataset at hand, a hybrid CNN-BiLSTM attention ensemble model is\nproposed that competently combines the intermediate sentence representations\ngenerated by the convolutional neural network and the bidirectional long\nshort-term memory, leading to efficient text classification. Experimental\nresults show that the proposed model outperforms various baseline machine\nlearning and deep learning models in the given task, giving the best validation\naccuracy of 89.57\\% and f1-score of 0.8875. Furthermore, the solution resulted\nin the best system submission for this subtask, giving a test accuracy of\n64.26\\% and f1-score of 0.6157, transcending the performances of other teams as\nwell as the baseline system given by the organizers of the shared task.\n
Paper
References (31)
Scroll for more · 19 remaining