Large Language Models (LLMs), such as GPT, have recently been deployed as agents in interactive environments, including games. While these models exhibit strong natural language understanding capabilities, they often produce incorrect or inconsistent decisions in strategic games that require strict rule adherence and multi-step planning. This project investigates the decision-making behavior of an LLM-based agent in the game of Tic-Tac-Toe. A simulator was developed to enable gameplay between the LLM and a human player, and multiple prompt designs were evaluated. The experiments compare prompts with no contextual guidance, basic rule descriptions, example-based contexts, and explicit strategic instructions. The results demonstrate that prompt design plays a critical role in improving the reliability of LLM-based game agents and offer insights into how decision-making accuracy can be significantly enhanced without modifying the underlying model.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex