Efficient Website Classification: Using Small-Scale Machine Learning Models for a Broad Range of Categories
With the ever-expanding amount of content on the internet, efficiently and accurately categorizing web pages has recently become a significant problem. Current web page classification models for categorizing web pages face many challenges: they either require large memory footprints, fail to maintain high accuracy, lack sufficient categories to cover the vast scope of web pages, or struggle to handle large datasets. To address these issues, this study proposes three small-scale fastText machine learning models with memory footprints of 154 MB, 47 MB, and 3 MB. These newly developed models use visible text and metadata from web pages and are trained on a dataset of over 2.5 million manually categorized English-language pages, sourced from Curlie, Common Crawl, and other internet sources. These pages were manually reviewed and categorized, marking this study the first to use 345 categories, providing a comprehensive representation of most web pages online. By applying quantization techniques, the models have achieved high accuracy with F1-scores of 90%, 91%, and 90%, respectively, despite their small memory sizes. This proposed approach provides an efficient solution for real-time, large-scale web page classification, minimizing computational resources while advancing scalable applications in content filtering, security, and information retrieval.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex