"You still have to study" -- On the Security of LLM generated code

We witness an increasing usage of AI-assistants even for routine (classroom) programming tasks. However, the code generated on basis of a so called"prompt"by the programmer does not always meet accepted security standards. On the one hand, this may be due to lack of best-practice examples in the training data. On the other hand, the actual quality of the programmers prompt appears to influence whether generated code contains weaknesses or not. In this paper we analyse 4 major LLMs with respect to the security of generated code. We do this on basis of a case study for the Python and Javascript language, using the MITRE CWE catalogue as the guiding security definition. Our results show that using different prompting techniques, some LLMs initially generate 65% code which is deemed insecure by a trained security engineer. On the other hand almost all analysed LLMs will eventually generate code being close to 100% secure with increasing manual guidance of a skilled engineer.

Paper

References (21)

112023. GitHub Copilot now has a better AI model and new capabilities2023 · , last access
122023. BestpracticesforpromptengineeringwithOpenAIAPI2023 · help.openai.com/en/articles/

Scroll for more · 9 remaining

Similar papers

© 2026 NYSGPT2525 LLC