OmniBench-RAG: A Multi-Domain Evaluation Platform for Retrieval-Augmented Generation Tools

While Retrieval Augmented Generation (RAG) is widely adopted to enhance LLMs, evaluating its true performance benefits in a reproducible and interpretable way remains challenging. Existing methods often fall short: they lack domain coverage, employ coarse metrics that miss sub-document precision, fail to capture computational trade-offs, and suffer from data leakage in some evaluation datasets (leading to unfair assessment results). Most critically, they provide no standardized framework for comparing RAG effectiveness across different models and domains.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC