Predicting Dewey Decimal Classification Notation from Book Titles: Machine Learning Approach

This study examines whether supervised machine learning can improve the Dewey Decimal Classification (DDC) in an academic library, utilising only book titles as input. A dataset of 6,987 monograph records with cataloguer-assigned DDC numbers was exported from the library catalogue and transformed into TF-IDF title vectors. Four multi-class classifiers (linear SVM, Multinomial Naïve Bayes, logistic Regression, and k-NN) were evaluated for predicting the three-digit DDC class. The linear SVM achieved the best performance, correctly assigning the broad class for about 71 % of 1,398 held-out test titles. Building on this model, a hierarchical architecture was implemented: the global SVM first predicts the three-digit class and section-specific SVMs then attempt to predict the full DDC notation in selected high-volume classes. The hierarchical system generated a whole DDC suggestion for approximately 80 % of test titles, with an exact full-code match in about 59 % of these cases and higher accuracies in the major engineering, computing and management sections. The findings indicate that supervised models cannot replace professional cataloguing. However, they can function as a practical decision-support tool.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC