This study evaluates the effectiveness of three prompting strategies: standard prompting, chain-of-thought (CoT) prompting, and informed CoT prompting on the performance of the GPT-J model in solving mathematical reasoning tasks from the GSM8K dataset.Using the full test set of 1,319 problems, we assess the model's performance through accuracy, F1 score, BLEU, and ROUGE metrics.The findings suggest that while providing relevant context, such as math topics, can modestly enhance performance, the gains are limited.This underscores the importance of carefully designing prompts in adaptive systems and indicates that additional strategies may be necessary to achieve practical utility in educational applications.
Paper
Full text
Enhancing Mathematical Reasoning in GPT-J Through Topic-Aware Prompt Engineering
OpenAlex · Machine Learning and Data Classification · 2025
Abstract
This study evaluates the effectiveness of three prompting strategies: standard prompting, chain-of-thought (CoT) prompting, and informed CoT prompting on the performance of the GPT-J model in solving mathematical reasoning tasks from the GSM8K dataset. Using the full test set of 1,319 problems, we assess the model’s performance through accuracy, F1 score, BLEU, and ROUGE metrics. The findings suggest that while providing relevant context, such as math topics, can modestly enhance performance, the gains are limited. This underscores the importance of carefully designing prompts in adaptive systems and indicates that additional strategies may be necessary to achieve practical utility in educational applications.