Document Clustering vs Topic Models: A Case Study

Document collections can be characterised in a variety of ways. Two key approaches are clustering, which partitions collections into subcollections with the expectation that the contents will be thematically linked, and topic models, which describe the contents in terms of weighted lists of words that are expected to represent different themes. In this paper, we report experiments on the observed relationship between clusters and topic models in a preliminary study of a large text collection. Both produce results that appear cohesive in their own right, but surprisingly – given the very different ways in which they are formed – the descriptions of the collections that they generate are strongly similar. This unexpected mutual reinforcement creates confidence in both approaches as tools for annotating and describing the contents of document collections.

Paper

Full text

PDF

Document Clustering vs Topic Models: A Case Study

Semantic Scholar · Computer Science · 2021

Abstract

Document collections can be characterised in a variety of ways. Two key approaches are clustering, which partitions collections into subcollections with the expectation that the contents will be thematically linked, and topic models, which describe the contents in terms of weighted lists of words that are expected to represent different themes. In this paper, we report experiments on the observed relationship between clusters and topic models in a preliminary study of a large text collection. Both produce results that appear cohesive in their own right, but surprisingly – given the very different ways in which they are formed – the descriptions of the collections that they generate are strongly similar. This unexpected mutual reinforcement creates confidence in both approaches as tools for annotating and describing the contents of document collections.

Similar papers

© 2026 NYSGPT2525 LLC