MultiVer: Zero-Shot Multi-Agent Vulnerability Detection

We present MultiVer, a zero-shot multi-agent system for vulnerability detection that achieves state-of-the-art recall without fine-tuning. A four-agent ensemble (security, correctness, performance, style) with union voting achieves 82.7% recall on PyVul, exceeding fine-tuned GPT-3.5 (81.3%) by 1.4 percentage points -- the first zeroshot system to surpass fine-tuned performance on this benchmark. On SecurityEval, the same architecture achieves 91.7% detection rate, matching specialized systems. The recall improvement comes at a precision cost: 48.8% precision versus 63.9% for fine-tuned baselines, yielding 61.4% F1. Ablation experiments isolate component contributions: the multi-agent ensemble adds 17 percentage points recall over single-agent security analysis. These results demonstrate that for security applications where false negatives are costlier than false positives, zero-shot multi-agent ensembles can match and exceed fine-tuned models on the metric that matters most.

Paper

References (16)

06Claude opus 4.5 model card2025 · Anthropic
07Multi-agent contextual reasoning for vulnerability detection2025 · MAVUL: Cross-function vulnerability tracking
08Retrieval-augmented vulnerability detection with knowledge-level analysis2024 · arXiv preprint
09Codeqwen1.5: An open-source code language model. Technical Report , 2024. 61.8% F1 zero-shot, 66.9% fine-tuned on PyVul
10When to vote, when to debate: Understanding multi-agent llm reasoningACL
11Codeql: Semantic code analysisGitHub
12Bandit: Security linter for pythonbandit.readthedocs.io/

Scroll for more · 4 remaining

Similar papers

© 2026 NYSGPT2525 LLC