Nowadays, a number of short messages or short text contents created on the Internet are rapidly increasing. Tasks to manipulate, analyze, and extract knowledge from them lead mining techniques such as text clustering to become more important. However, applying traditional text clustering algorithms which consider only common words or phrases to group short texts is inefficient due to the problem of sparsity. In this paper, we propose a new clustering technique, called word semantic graph clustering, based on the use of text concepts. We apply the word embedding model from Word2Vec to capture the semantic meaning of words and later construct semantic subgraphs in which those words represented as vertices are connected by some high semantic similarities. Finally, short text documents will be assigned to the same cluster if they contain at least one word belonging to the same semantic subgraph. Experimental results conducted on two real datasets show that the proposed approach outperforms the state-of-the-art text clustering algorithms. In addition, it can also produce more appropriate label for each cluster than the comparative algorithms do.
Paper
Full text
Short Text Clustering Based on Word Semantic Graph with Word Embedding Model
Semantic Scholar · Computer Science · 2018
Abstract
Nowadays, a number of short messages or short text contents created on the Internet are rapidly increasing. Tasks to manipulate, analyze, and extract knowledge from them lead mining techniques such as text clustering to become more important. However, applying traditional text clustering algorithms which consider only common words or phrases to group short texts is inefficient due to the problem of sparsity. In this paper, we propose a new clustering technique, called word semantic graph clustering, based on the use of text concepts. We apply the word embedding model from Word2Vec to capture the semantic meaning of words and later construct semantic subgraphs in which those words represented as vertices are connected by some high semantic similarities. Finally, short text documents will be assigned to the same cluster if they contain at least one word belonging to the same semantic subgraph. Experimental results conducted on two real datasets show that the proposed approach outperforms the state-of-the-art text clustering algorithms. In addition, it can also produce more appropriate label for each cluster than the comparative algorithms do.