Keyword Extraction for Educational Content Generation Using NLP Algorithms

Automated keyword extraction plays a crucial role in processing and organizing textual data, particularly in the context of academic course syllabuses. This research aims to evaluate and compare three widely used keyword extraction algorithms such as KeyBERT, TF-IDF, and YAKE to determine their precision, reliability, and relevance in extracting meaningful keywords from course descriptions, objectives, and weekly topics. These three methods were selected based on distinct characteristics: TF-IDF as a statistical approach, YAKE as an unsupervised method relying on text features, and KeyBERT as an embedding-based model leveraging contextual similarity. By applying these algorithms to a dataset of 255 course syllabuses, we analyzed their performance using key metrics such as precision, recall, F1-score, and contextual relevance. The evaluation results indicate that YAKE demonstrated the highest precision, making it more suitable for extracting domain-relevant keywords. TF-IDF achieved the highest recall, capturing a broader set of terms, while KeyBERT balanced both precision and recall effectively. These findings highlight the strengths and limitations of each method, providing insights into their applicability for academic content analysis. The study contributes to the ongoing development of automated keyword extraction techniques by offering a comparative assessment that can inform future improvements and applications in educational and research contexts.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC