Validating MCP-Based, GenAI-Powered Automated Essay Scoring and Feedback System in English Language Assessment

In recent years, generative artificial intelligence (GenAI), such as Claude and ChatGPT, are being used to develop automated scoring systems in various contexts of English language assessment. Oftentimes, the model context protocol (MCP) is employed as a platform for coordinating between LLM and AI agents in the AI agent (or agentic AI) systems. The main goals of the current paper are to: (a) build MCP-based, GenAI-powered automated essay scoring models (zero-shot, few-shot) by using Claude as the base GenAI program, (b) use these scoring models to evaluate the EFL learners’ essays enrolled in a college English course at a university located in Seoul, South Korea, and (c) investigate not only the reliability and validity of automated essay scores in comparison with human rater scores but also the usefulness of generative feedback from the automated scoring system. Results of analyses show that relatively high score agreement can be achieved between the Claude-powered automated essay scoring system and human raters; high correlations are observed between the automated and human rater scores; and the automated scoring systems also produced very useful diagnostic feedback for students. Implications of the findings are also discussed along with some recommendations for future avenues for research on GenAI-based automated essay scoring in South Korea.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC