Prompt to Protection: A Comparative Study of Multimodal LLMs in Construction Hazard Recognition
The recent emergence of multimodal large language models (LLMs) has introduced new opportunities for improving visual hazard recognition on construction sites. However, despite growing interest in their applications, there has been limited investigation into how different LLMs perform in safety-critical visual tasks within the construction domain. To address this gap, this study conducts a comparative evaluation of five state-of-the-art LLMs: GPT-4o, GPT-5, GPT-4.1, Claude 4.1 Opus, and Gemini 2.5 Pro, to assess their ability to identify potential hazards from real-world construction images. Each model was tested under three prompting strategies: zero-shot, few-shot, and chain-of-thought (CoT). Quantitative analysis was performed using precision, recall, and F1-score metrics across all conditions. Results reveal that prompting strategy significantly influenced hazard recognition performance. CoT prompting consistently produced the highest accuracy across models, with GPT-5 and GPT-4.1 achieving superior scores in most settings. The results suggest that structured prompt design can elevate LLM performance to significant levels, offering a cost-effective pathway for small firms and contractors to improve hazard recognition. Furthermore, LLM outputs can be repurposed as explanatory content for safety training and toolbox talks, supporting long-term safety culture development.