TAA-EPLMR: Threat Actor Attribution via Evidence Path-Enhanced Large Language Model Reasoning
Threat actor attribution (TAA) is a complex task that requires multi-source intelligence fusion and semantic reasoning. In cyber threat intelligence (CTI) sharing, indicators of compromise (IOCs), with their diverse types and interconnections, provide critical evidence chains for TAA. However, existing methods primarily rely on small-scale intelligence data and embedding models, thereby limiting performance. Large language models (LLMs), with advanced semantic understanding and in-context learning capabilities, provide a promising approach to the complex semantic reasoning challenge in TAA. In this paper, we propose TAA-EPLMR, an evidence path-enhanced LLM reasoning approach that introduces a novel paradigm for TAA, cohesively integrating CTI knowledge graphs (CTIKGs) with large language models. We first define multi-level evidence path patterns (EPPs) grounded in CTI-based attribution semantics. We leverage these EPPs to retrieve candidate evidence paths from the CTI-KG, apply an attacker-discriminability-based pruning algorithm, and perform attacker-wise path aggregation to obtain refined evidence subgraphs for the candidate attackers. Furthermore, we design a chain of thought grounded in evidenceaware attribution logic and progressively challenging few-shot demonstrations. We prompt the LLM to infer threat actor attribution using the above information and generate attribution explanations along with confidence scores. Experiments on three datasets with varying completeness and noise levels consistently show that TAA-EPLMR outperforms all baselines and enhances the explainability and credibility of attribution reasoning.
Paper
Full text
TAA-EPLMR: Threat Actor Attribution via Evidence Path-Enhanced Large Language Model Reasoning
Semantic Scholar · Computer Science · 2025
Abstract
Threat actor attribution (TAA) is a complex task that requires multi-source intelligence fusion and semantic reasoning. In cyber threat intelligence (CTI) sharing, indicators of compromise (IOCs), with their diverse types and interconnections, provide critical evidence chains for TAA. However, existing methods primarily rely on small-scale intelligence data and embedding models, thereby limiting performance. Large language models (LLMs), with advanced semantic understanding and in-context learning capabilities, provide a promising approach to the complex semantic reasoning challenge in TAA. In this paper, we propose TAA-EPLMR, an evidence path-enhanced LLM reasoning approach that introduces a novel paradigm for TAA, cohesively integrating CTI knowledge graphs (CTIKGs) with large language models. We first define multi-level evidence path patterns (EPPs) grounded in CTI-based attribution semantics. We leverage these EPPs to retrieve candidate evidence paths from the CTI-KG, apply an attacker-discriminability-based pruning algorithm, and perform attacker-wise path aggregation to obtain refined evidence subgraphs for the candidate attackers. Furthermore, we design a chain of thought grounded in evidenceaware attribution logic and progressively challenging few-shot demonstrations. We prompt the LLM to infer threat actor attribution using the above information and generate attribution explanations along with confidence scores. Experiments on three datasets with varying completeness and noise levels consistently show that TAA-EPLMR outperforms all baselines and enhances the explainability and credibility of attribution reasoning.