Web Document Categorization Using Naive Bayes Classifier and Latent Semantic Analysis

A rapid growth of web documents due to heavy use of World Wide Web\nnecessitates efficient techniques to efficiently classify the document on the\nweb. It is thus produced High volumes of data per second with high diversity.\nAutomatically classification of these growing amounts of web document is One of\nthe biggest challenges facing us today. Probabilistic classification algorithms\nsuch as Naive Bayes have become commonly used for web document classification.\nThis problem is mainly because of the irrelatively high classification accuracy\non plenty application areas as well as their lack of support to handle high\ndimensional and sparse data which is the exclusive characteristics of textual\ndata representation. also it is common to Lack of attention and support the\nsemantic relation between words using traditional feature selection method When\ndealing with the big data and large-scale web documents. In order to solve the\nproblem, we proposed a method for web document classification that uses LSA to\nincrease similarity of documents under the same class and improve the\nclassification precision. Using this approach, we designed a faster and much\naccurate classifier for Web Documents. Experimental results have shown that\nusing the mentioned preprocessing can improve accuracy and speed of Naive Bayes\navailably, the precision and recall metrics have indicated the improvement.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC