Stance classification of user responses plays a key role in rumor detection on social media platforms. Stances are commonly divided into four categories: support, deny, query and comment, where the first three are particularly important for determining the confidence of rumors. Since people seldom express definite stance under rumors with unclear authenticity, it is difficult to judge such stances. Machine learning approaches have been proposed to address this problem with either manually designed features or automatically extracted features. However, the problem is related to many aspects of attributes and it is somehow difficult to summarize a comprehensive feature template or to learn an effective feature representation. In this work, we conduct an in-depth study on the feature engineering for this task. We screen out 18 salient features in three aspects including text content, user portrait and propagation state. The experimental results on the RumorEval dataset show that coupled with these 18 features, a traditional logistic regression classifier even achieves the state-of-the-art performance and outperforms some complex neural networks such as long-short term memory networks that uses the same feature template or automatic feature extraction.
Paper
Full text
Rumor Stance Classification via Machine Learning with Text, User and Propagation Features
Semantic Scholar · Computer Science · 2019
Abstract
Stance classification of user responses plays a key role in rumor detection on social media platforms. Stances are commonly divided into four categories: support, deny, query and comment, where the first three are particularly important for determining the confidence of rumors. Since people seldom express definite stance under rumors with unclear authenticity, it is difficult to judge such stances. Machine learning approaches have been proposed to address this problem with either manually designed features or automatically extracted features. However, the problem is related to many aspects of attributes and it is somehow difficult to summarize a comprehensive feature template or to learn an effective feature representation. In this work, we conduct an in-depth study on the feature engineering for this task. We screen out 18 salient features in three aspects including text content, user portrait and propagation state. The experimental results on the RumorEval dataset show that coupled with these 18 features, a traditional logistic regression classifier even achieves the state-of-the-art performance and outperforms some complex neural networks such as long-short term memory networks that uses the same feature template or automatic feature extraction.