Large language models (LLMs) are playing an increasingly pivotal role in LegalAI. However, existing benchmarks are primarily tailored for legal professionals, emphasizing deep reasoning and explainability. While public-facing legal applications demand outputs that are direct, actionable, and accessible, a need largely over-looked by current evaluation frameworks. To bridge this gap, we propose a public-oriented LegalAI benchmark grounded in legal func-tionalism and genre analysis. Specifically, we categorize public legal demands into two core tasks: Instant Question Answering and Legal Text Generation . We further introduce three public-oriented evaluation dimensions: legal normativity , content relevance , and format usability , which collectively assess the practical validity and user readiness of model outputs. To reflect real-world lay user usage, we evaluate 17 LLMs on Pub-LawBench using only simple prompts and Chain-of-Thought under a vanilla inference setting, excluding complex techniques like RAG or agent-based methods inaccessible to non-experts. Experiments reveal limitations of current LLMs in delivering effective public-oriented legal assistance, high-lighting the need for more user-centric model development
Paper
Full text
Pub-LawBench: Public-Oriented Benchmarking for LegalAI
Semantic Scholar · 2026
Abstract
Large language models (LLMs) are playing an increasingly pivotal role in LegalAI. However, existing benchmarks are primarily tailored for legal professionals, emphasizing deep reasoning and explainability. While public-facing legal applications demand outputs that are direct, actionable, and accessible, a need largely over-looked by current evaluation frameworks. To bridge this gap, we propose a public-oriented LegalAI benchmark grounded in legal func-tionalism and genre analysis. Specifically, we categorize public legal demands into two core tasks: Instant Question Answering and Legal Text Generation . We further introduce three public-oriented evaluation dimensions: legal normativity , content relevance , and format usability , which collectively assess the practical validity and user readiness of model outputs. To reflect real-world lay user usage, we evaluate 17 LLMs on Pub-LawBench using only simple prompts and Chain-of-Thought under a vanilla inference setting, excluding complex techniques like RAG or agent-based methods inaccessible to non-experts. Experiments reveal limitations of current LLMs in delivering effective public-oriented legal assistance, high-lighting the need for more user-centric model development
References (70)
Scroll for more · 38 remaining