Crimson: Empowering Strategic Reasoning in Cybersecurity through Large Language Models

We present Crimson, a system designed to bolster the strategic reasoning skills of Large Language Models (LLMs) within the realm of cybersecurity. By linking CVEs to MITRE ATT&CK techniques, Crimson enhances threat forecasting and strategic defense measures. Our methodology includes the definition and assessment of cybersecurity strategic tasks, utilizing a thorough human-in-the-loop data synthesis process to create the CVE-to-ATT&CK Mapping (CVEM) dataset. Additionally, we improve the reasoning capabilities of LLMs through an innovative Retrieval-Aware Training (RAT) technique and its superior version, RAT-R.Our research reveals that a language model with 7 billion parameters, optimized using our methods, closely matches the performance of GPT-4. It significantly reduces errors and hallucinations while outperforming other models in strategic reasoning. Additionally, tailoring embedding models for specific domains greatly improves cybersecurity performance, highlighting the effectiveness of our approaches. Utilizing Crimson to transform raw vulnerability data into structured, actionable insights enhances proactive cybersecurity measures.

Paper

Similar papers

© 2026 NYSGPT2525 LLC