LLM-Driven Usefulness Judgment for Web Search Evaluation

Evaluation is fundamental to optimizing search experiences and supporting diverse user intents in Information Retrieval (IR). Traditional search evaluation methods primarily rely on relevance labels, which assess how well retrieved documents match a user's query. However, relevance alone fails to capture a search system's effectiveness in helping users achieve their goals, making usefulness a critical evaluation criterion. Recent LLM-enabled evaluation works have mostly focused on relevance label generation. In this paper, we explore an alternative approach: LLM-generated usefulness labels that incorporate implicit and explicit user behavior signals along with relevance. We introduce Task-aware Rubric-based Usefulness Evaluation (TRUE), a reproducible rubric-driven framework that leverages iterative sampling and Chain-of-Thought reasoning to model complex search behavior patterns. Our study shows that: (i) pre-trained LLMs can generate moderate usefulness labels with rich session-level context; and (ii) LLMs with TRUE outperform state-of-the-art methods and our systematically constructed baseline. We further analyze the relationship between usefulness and user satisfaction by comparing label generation with and without satisfaction signals, quantifying their correlation and impact on model behavior. Additionally, we conduct an ablation study to identify key features for accurate usefulness label generation, enabling cost-effective evaluation. Overall, this work advances LLM-based evaluation beyond relevance by proposing a reproducible and scalable framework for usefulness judgment, addressing key reproducibility challenges.

Paper

References (55)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC