Revealing Learner Interests through Topic Mining from Question-Answering Data

In a question-answering system, learner generated content including asked and answered questions is a meaningful resource to capture learning interests. This paper proposes an approach based on question topic mining for revealing learners' concerned topics in real community question-answering systems. The authors' approach firstly preprocesses all questions associated with learners. Afterwards, it analyzes each question with text features and generates a weight feature matrix using a revised TF/IDF method. In order to decrease the sparsity issue of data distribution, the authors employ three concept-mapping strategies including named entity recognition, synonym extension, and hyponym replacement. Applying an SVM classifier, their approach categorizes user questions into representative topics. Three experiments are conducted based on a TREC dataset and an actual dataset containing 1,120 questions posted by learners from a commercial question-answering community. Results demonstrate the effectiveness of the method compared with conventional classifiers as baselines.

Paper

Full text

PDF

Revealing Learner Interests through Topic Mining from Question-Answering Data

Semantic Scholar · Computer Science · 2017

Abstract

In a question-answering system, learner generated content including asked and answered questions is a meaningful resource to capture learning interests. This paper proposes an approach based on question topic mining for revealing learners' concerned topics in real community question-answering systems. The authors' approach firstly preprocesses all questions associated with learners. Afterwards, it analyzes each question with text features and generates a weight feature matrix using a revised TF/IDF method. In order to decrease the sparsity issue of data distribution, the authors employ three concept-mapping strategies including named entity recognition, synonym extension, and hyponym replacement. Applying an SVM classifier, their approach categorizes user questions into representative topics. Three experiments are conducted based on a TREC dataset and an actual dataset containing 1,120 questions posted by learners from a commercial question-answering community. Results demonstrate the effectiveness of the method compared with conventional classifiers as baselines.

Similar papers

© 2026 NYSGPT2525 LLC