The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelines
Cybersecurity operations are increasingly adopting agentic AI solutions due to the time-critical and complex decision-making in security operations centers (SOCs). While large language models (LLMs) are good with summarization tasks or interpreting structured and unstructured reports, real-world SOC workflows have additional requirements such as access to original logs, reproducibility and accountability to triage security incidents. For example, analysts routinely correlate alerts to understand the kill-chain of the cyber-attack and analyze the event telemetries to identify the root cause event which may not have triggered an alert. Incorrect and incomplete automations in such settings can directly impact production systems and business operations.In this survey, we examine the architectural shifts from single-model assistants to tool-augmented agents, distributed multiagent systems, and schema-constrained investigation pipelines. We introduce a five-generation taxonomy that represents the evolution of agentic AI systems, their limitations and risks across different parameters, such as reasoning depth, tool interaction, memory, reproducibility and safety. We also review the emerging benchmarks to evaluate cyber-oriented agents and identify open challenges including response validation, tool-use correctness, multi-agent coordination, long-horizon reasoning and safeguards for high-impact actions. Finally, we discuss how these challenges influence deployment decisions in operational SOC environments. Our analysis provides a structured perspective on the current state of agentic AI in cybersecurity and highlights the technical and governance considerations necessary for its deployment.
Paper
References (54)
Scroll for more · 38 remaining