The LLM Effect on IR Benchmarks: A Meta-Analysis of Effectiveness, Baselines, and Contamination

Benchmark collections have long enabled controlled comparison and cumulative progress in Information Retrieval (IR). However, prior meta-analyses show that reported effectiveness gains often fail to accumulate, in part due to weak or outdated baselines. Large language models (LLMs) are increasingly used in retrieval pipelines, yet their impact on established IR benchmarks has not been systematically analyzed. We analyze 179 publications reporting on the TREC Robust04 collection and the TREC Deep Learning 2020 (DL20) Passage Retrieval benchmark, using ACM Digital Library keyword search supplemented by citation-graph backtracking for Robust04. We observe what we term an LLM effect: recent systems incorporating LLM components achieve 8.8% higher nDCG@10 on DL20 than the best TREC 2020 result and 11.9% higher on Robust04 than the strongest pre-2024 result. However, evaluation practice has shifted from MAP to nDCG@10 over the same window, and our adaptation of the Data Contamination Quiz reveals 12-41% contamination across two widely-used LLM rerankers. Filtering contaminated topics shows no statistically significant effectiveness difference, but small samples and the uncertainty of adapting contamination detection to reranking prevent us from ruling out memorization as a contributing factor. We read the LLM effect as real but unverified: visible in the aggregate numbers, but not cleanly separable from metric drift or pretraining overlap.

Paper

References (21)

Scroll for more · 9 remaining

Similar papers

© 2026 NYSGPT2525 LLC