On Technique Identification and Threat-Actor Attribution using LLMs and Embedding Models

Attribution of cyber-attacks remains a complex but critical challenge for defenders. Manual extraction of behavioral indicators from dense forensic documentation causes significant attribution delays, especially following major incidents at the international scale. This research evaluates large language models (LLMs) for cyber-attack attribution based on behavioral indicators extracted from forensic documentation. We test OpenAI's GPT-4 and text-embedding-3-large for identifying threat actors' tactics, techniques, and procedures (TTPs) by comparing LLM-generated TTPs against human-generated data from MITRE ATT&CK datasets. Our framework identifies TTPs from text using vector embedding search and builds profiles to attribute new attacks. Key contributions include: (1) assessing off-the-shelf LLMs for TTP extraction and attribution, and (2) developing an end-to-end pipeline from raw documentation to threat-actor prediction. This research finds that standard LLMs generate noisy TTP datasets, resulting in low similarity to human-generated datasets. However, the TTPs generated maintain frequency patterns aligned with MITRE datasets and enable attribution performance that outperforms a random guessing baseline. Project code available at: https://github.com/kylag/ttp_attribution.

Paper

References (34)

Scroll for more · 22 remaining

Similar papers

© 2026 NYSGPT2525 LLC