Syntagmatic Distractor Generation for Multiple-Choice Language Tests: A Large Language Model-Based Approach
High-quality distractors are crucial for multiple-choice vocabulary questions in language proficiency tests to ensure validity and reduce guessing. This study proposes an automatic distractor generation system for TOEFL vocabulary items using a syntagmatic approach with the LLaMA 3 Large Language Model. The system generates distractors by leveraging the target word's sentential context, followed by filtering using semantic embeddings and collocational embeddings to ensure contextual relevance and semantic distinction. Evaluation on 21 TOEFL vocabulary items produced 63 distractors, with 49.21% classified as contextually appropriate and sufficiently distinct, while 41.27% were irrelevant. These results highlight the potential and challenges of using LLMs for syntagmatic distractor generation. Future work should focus on enhanced filtering methods and expert validation to improve pedagogical quality.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex