Security Vulnerability Patterns in AI-Generated Code: A Cross-Model Comparative Study

LLM-based coding tools enable non-expert users to generate routine automation scripts that may enter enterprise workflows without meaningful security review. This study examines that risk directly. Code was collected from ChatGPT, Microsoft Copilot, and Google Gemini using identical prompts across three automation domains. Claude Code performed a standardized vulnerability review. Each identified vulnerability was scored using CVSS v3.1 and mapped to the OWASP Top 10:2021 and the MITRE ATT&CK frameworks. Every script contained exploitable vulnerabilities. Nine of the 17 identified vulnerability classes appeared in code from all three models, while 14 of the 17 vulnerability classes appeared in at least two models. The weighted CVSS scores across platforms differed by less than 10%. The risk is not tied to any particular model but rather to the task category. Organizations should therefore ask not which tool to trust, but instead whether LLM-generated automation code should be deployed without review.

Paper

References (17)

06Anthropic, Claude 3.7 sonnet , Feb2025 · anthropic.com/news
07Anthropic, Enabling Claude Code to work more autonomously2025 · anthropic.com/news
08“Common vulnerability scoring system v3.1: Specification document,”2019 · Forum
12“Cost of a data breach report 2025,”

Scroll for more · 5 remaining

Similar papers

© 2026 NYSGPT2525 LLC