Toward Standardized Interactivity: An AI-Enabled Adaptive Speaking Task for Computer-Delivered Large-Scale Assessment

Despite task interactivity eliciting different aspects of language use, more interactive, adaptive speaking tasks have been challenging to implement in computer-delivered large-scale high-stakes contexts. Recent advances in large language models (LLMs) offer a potential means to address this tension through spoken dialogue systems (SDSs) but fall short of full adaptivity. To enhance adaptivity in an LLM-enhanced SDS-delivered interview task, this article proposes two solutions: (1) LLM-driven response evaluation using question-specific task completion rubrics for adaptive follow-up question selection and (2) item response theory (IRT) scoring accounting for prompt variability. We examined the extent to which these solutions functioned to simulate adaptivity, from 5,909 test takers’ responses to an adaptive speaking task with an avatar and a non-adaptive monologic task. Automated task completion evaluations were comparable to expert ratings, though follow-up question evaluation showed low agreement both among human raters and between raters and the system. IRT scoring yielded higher test–retest reliability, and the adaptive task elicited more reciprocal language use than the monologic task. Unlike earlier SDSs that prioritized consistency over adaptivity, the real-time meaning-level evaluation with IRT modeling strikes a balance between standardization and adaptivity. These findings support the viability of adaptive speaking tasks in computer-based standardized assessments.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC