Categorisation of text is significant trend that ultimately appears owing to the internet revolution nowadays resulting in enormous amounts data that depend on various languages. The Arabic language is one of the most commonly used languages all over the world; it is considered the fifth most spoken one. Various challenges occur through processing and classifying of Arabic text since it has more sophisticated techniques than the English language. These challenges are clear owing to the Arabic language's variation in shape, structure and component; besides, there is a lack of adequate studies discussing Arabic text classification. This research seeks to form a general point of view by categorising different techniques of Arabic text classification for helping new researches concerning this domain. Also, it shows some of prior information and innovative designs about Arabic text classification. Besides, it mentions various works that have discussed classifying Arabic text, with regard to data sets, categories and pre-processes steps, classification mechanism and assessment procedure for those techniques. These discussions aim to conclude a comprehensive overview through forming a general framework for all researchers about this domain via examining the defects of the prior studies, and then the possibility of presenting more advanced directions.
Paper
Full text
A survey of Arabic text classification approaches
Semantic Scholar · Computer Science · 2019
Abstract
Categorisation of text is significant trend that ultimately appears owing to the internet revolution nowadays resulting in enormous amounts data that depend on various languages. The Arabic language is one of the most commonly used languages all over the world; it is considered the fifth most spoken one. Various challenges occur through processing and classifying of Arabic text since it has more sophisticated techniques than the English language. These challenges are clear owing to the Arabic language's variation in shape, structure and component; besides, there is a lack of adequate studies discussing Arabic text classification. This research seeks to form a general point of view by categorising different techniques of Arabic text classification for helping new researches concerning this domain. Also, it shows some of prior information and innovative designs about Arabic text classification. Besides, it mentions various works that have discussed classifying Arabic text, with regard to data sets, categories and pre-processes steps, classification mechanism and assessment procedure for those techniques. These discussions aim to conclude a comprehensive overview through forming a general framework for all researchers about this domain via examining the defects of the prior studies, and then the possibility of presenting more advanced directions.