Design and Development of a Human-Evaluation Platform for Large Language Model-Generated Datasets

This study aims to design and develop a human evaluation platform for assessing educational question-answer (QA) datasets generated by closed-source large language models (LLMs). As the use of generative AI expands in educational settings, there is a growing need for structured review systems capable of verifying the reliability and curricular alignment of high-quality datasets. The proposed platform enables experts to evaluate QA datasets by uploading data, ranking multiple responses, and providing qualitative feedback. To support this process, a web-based system was implemented, incorporating features such as metadata management, expert account creation and assignment, and rank-based evaluation. By introducing a practical tool that supplements human involvement in dataset validation, this study lays the groundwork for the evaluation and refinement of educational language models.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC