Large Language Models for Software Testing Education: an Experience Report

Large Language Models (LLMs) are rapidly entering software engineering practice, including software testing. As students increasingly use LLMs in testing tasks, software testing education must adapt. Yet we still know little about how such use shapes students' testing behaviors, judgment, and learning. Systematic investigation is therefore needed. This paper reports a mixed-methods, two-phase exploratory study of human-LLM collaboration in software testing education. In Phase I, we analyze classroom artifacts and interaction records from 15 students together with a national software testing competition survey (337 valid responses) to identify recurring prompt-related difficulties across testing tasks. We find missing context, insufficient constraints, rigid one-shot prompting, and limited strategy-driven iteration. Automated test script generation emerges as the most heterogeneous and effort-intensive context. In Phase II, we translate these findings into a lightweight, stage-aware prompt scaffold for test script generation and report descriptive shifts in students' articulation of execution-relevant information. From a learning-behavior perspective, this paper characterizes task-dependent difficulties and interaction patterns in students' use of LLMs for software testing. It distinguishes relatively stable interaction contexts from more heterogeneous ones and shows how observed difficulties can be translated into classroom support. Overall, the findings provide an empirical foundation for learning-oriented integration of LLMs into software testing education.

Paper

Similar papers

© 2026 NYSGPT2525 LLC