Large Language Models for Secure Code Assessment: A Multi-Language Empirical Study

Most vulnerability detection studies rely on datasets tied to specific programming languages and within-project settings, limiting language diversity and the assessment of model generalizability. Consequently, the effectiveness of deep learning methods, including large language models (LLMs), beyond such restricted scenarios remains largely unexplored. In this paper, we evaluate six state-of-the-art LLMs (GPT-3.5-Turbo, GPT-4 Turbo, GPT-4o, CodeLlama-7B, CodeLlama-13B, Gemini 1.5 Pro) on vulnerability detection and CWE classification across five languages (Python, C, C++, Java, JavaScript). We compiled a multi-language dataset from diverse sources and investigated different prompt and role strategies. Our results show that GPT-4o achieves the best performance, particularly in few-shot settings. To examine practical applicability, we developed CodeGuardian, a VSCode extension that enables real-time LLM-assisted vulnerability detection. In a user study with 22 professional developers, CodeGuardian doubled accuracy and halved task completion time compared to a control group.

Paper

References (64)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC