Large Language Models (LLMs) have shown significant challenges in detecting and repairing vulnerable code, particularly when dealing with vulnerabilities involving multiple aspects, such as variables, code flows, and code structures. In this study, we utilize GitHub Copilot as the LLM and focus on buffer overflow vulnerabilities. Our experiments reveal a notable gap in GitHub Copilot's vulnerability repair abilities, with a 76% vulnerability detection rate but only a 15% vulnerability repair rate. To address this issue, we propose a context-aware prompt tuning technique to enhance Copilot's performance in repairing buffer overflow. By injecting a sequence of domain knowledge about the vulnerability, including various security and code contexts, we demonstrate that Copilot's vulnerability repair rate increases to 63%, representing more than four times the improvement compared to repairs without domain knowledge.
Paper
References (34)
Scroll for more · 22 remaining