Assisting the Grading of a Handwritten General Chemistry Exam with Artificial Intelligence

We explore the effectiveness and reliability of an artificial intelligence (AI)-based grading system for a handwritten general chemistry exam, comparing AI-assigned scores to human grading across various types of questions. Exam pages and grading rubrics were uploaded as images to account for chemical reaction equations, short and long open-ended answers, numerical and symbolic answer derivations, drawing, and sketching in pencil-and-paper format. Using descriptive linear-fit comparisons and psychometric evaluations, the investigation reveals high agreement between AI and human graders for textual and chemical reaction questions, while highlighting lower reliability for numerical and graphical tasks. The findings emphasize the necessity for human oversight to ensure grading accuracy, based on selective filtering. The results indicate promising applications for AI in routine assessment tasks, though careful consideration must be given to student perceptions of fairness and trust in integrating AI-based grading into educational practice.

Paper

References (20)

03OpenAI, Reasoning models2025 · platform.openai.com/ docs/guides/reasoning
04Proceedings of the 21st International Conference on Artificial Intelligence in Education (AIED 2020)2020 · Cham
05Numerical: Students need to calculate values
06Lin’s CCC (level agreement): Values near 1 mean runs not only track each other strongly but also show negligible
07Observed Score ( s ij
08Expected Score
09S w is the typical run-to-run SD for the same target
10Fig. 6 An example of a false negative (problem 6-C-b)
11Azure AI Services/
12Multiple choice: Students need to select between different options

Scroll for more · 8 remaining

Similar papers

© 2026 NYSGPT2525 LLC