Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts

The safe deployment of large language models (LLMs) in high-stakes fields like biomedicine, requires them to be able to reason about cause and effect. We investigate this ability by testing 13 open-source LLMs on a fundamental task: pairwise causal discovery (PCD) from text. Our benchmark, using 12 diverse datasets, evaluates two core skills: 1) Causal Detection (identifying if a text contains a causal link) and 2) Causal Extraction (pulling out the exact cause and effect phrases). We tested various prompting methods, from simple instructions (zero-shot) to more complex strategies like Chain-of-Thought (CoT) and Few-shot In-Context Learning (FICL). The results show major deficiencies in current models. The best model for detection, DeepSeek-R1-Distill-Llama-70B, only achieved a mean score of $\mathbf{4 9. 5 7 \%}\left(C_{\text{detect}}\right)$, while the best for extraction, Qwen2.5-Coder-32B-Instruct, reached just 47.12 % ($C_{\text{extract}}$). Models performed best on simple, explicit, singlesentence relations. However, performance plummeted for more difficult (and realistic) cases, such as implicit relationships, links spanning multiple sentences, and texts containing multiple causal pairs. We provide a unified evaluation framework, built on a dataset validated with high inter-annotator agreement ($\kappa \geq 0.758$), and make all our data, code, and prompts publicly available to spur further research. Code available here: https://github.com/sydneyanuyah/CausalDiscovery

Paper

References (43)

Scroll for more · 31 remaining

Similar papers

© 2026 NYSGPT2525 LLC