AI-Powered Automated Penetration Testing in Kali Linux: An Enterprise-Scale Offensive Security Framework Driven by Reinforcement Learning and Large Language Models

: Offensive cybersecurity is changing because of AI-enabled automation, especially in penetration testing platforms like Kali Linux. There is little academic research evaluating the comparative performance, operational risks, and governance requirements of various tools, such as multi-agent testers like BreachSeek, Large Language Model (LLM) systems like PentestGPT and RapidPen, and Reinforcement Learning (RL) agents like DeepExploit. In order to close this gap, this study presents unique tests carried out on a simulated 20-node business environment and suggests a unified analytical methodology for assessing AI-assisted penetration testing. We assess RL-based, LLM-based, and hybrid multi-agent systems in terms of accuracy, exploitation speed, false-positive rates, failure modes, and adversarial brittleness using standardized attack chains. Results indicate that LLM-driven orchestrators perform better than RL agents in reconnaissance and decision-making (avg. 4.3 min to shell vs. 7.8 min), although RL agents suffer from sparse-reward inefficiencies and provide more consistent payload selection. In line with the EU AI Act, ISO/IEC 42001, and the NIST AI Risk Management Framework (AI RMF), the study incorporates a structured ethical and governance analysis that highlights vulnerabilities including dual-use misuse pathways, over-automation, and hallucinated exploit paths. This work promotes a research-based framework for safe, responsible adoption in cybersecurity operations and offers a fair, fact-based assessment of AI-powered penetration testing.

Paper

Full text

PDF

AI-Powered Automated Penetration Testing in Kali Linux: An Enterprise-Scale Offensive Security Framework Driven by Reinforcement Learning and Large Language Models

Semantic Scholar · 2026

Abstract

: Offensive cybersecurity is changing because of AI-enabled automation, especially in penetration testing platforms like Kali Linux. There is little academic research evaluating the comparative performance, operational risks, and governance requirements of various tools, such as multi-agent testers like BreachSeek, Large Language Model (LLM) systems like PentestGPT and RapidPen, and Reinforcement Learning (RL) agents like DeepExploit. In order to close this gap, this study presents unique tests carried out on a simulated 20-node business environment and suggests a unified analytical methodology for assessing AI-assisted penetration testing. We assess RL-based, LLM-based, and hybrid multi-agent systems in terms of accuracy, exploitation speed, false-positive rates, failure modes, and adversarial brittleness using standardized attack chains. Results indicate that LLM-driven orchestrators perform better than RL agents in reconnaissance and decision-making (avg. 4.3 min to shell vs. 7.8 min), although RL agents suffer from sparse-reward inefficiencies and provide more consistent payload selection. In line with the EU AI Act, ISO/IEC 42001, and the NIST AI Risk Management Framework (AI RMF), the study incorporates a structured ethical and governance analysis that highlights vulnerabilities including dual-use misuse pathways, over-automation, and hallucinated exploit paths. This work promotes a research-based framework for safe, responsible adoption in cybersecurity operations and offers a fair, fact-based assessment of AI-powered penetration testing.

References (11)

11ISO/IEC 42001:2023—Information technology-Artificial intelligence-Management system2023

Similar papers

© 2026 NYSGPT2525 LLC