NewsBench: A Systematic Evaluation Framework for Assessing Editorial Capabilities of Large Language Models in Chinese Journalism

We present NewsBench, a novel evaluation framework to systematically assess the capabilities of Large Language Models (LLMs) for editorial capabilities in Chinese journalism. Our constructed benchmark dataset is focused on four facets of writing proficiency and six facets of safety adherence, and it comprises manually and carefully designed 1,267 test samples in the types of multiple choice questions and short answer questions for five editorial tasks in 24 news domains. To measure performances, we propose different GPT-4 based automatic evaluation protocols to assess LLM generations for short answer questions in terms of writing proficiency and safety adherence, and both are validated by the high correlations with human evaluations. Based on the systematic evaluation framework, we conduct a comprehensive analysis of ten popular LLMs which can handle Chinese. The experimental results highlight GPT-4 and ERNIE Bot as top performers, yet reveal a relative deficiency in journalistic safety adherence in creative writing tasks. Our findings also underscore the need for enhanced ethical guidance in machine-generated journalistic content, marking a step forward in aligning LLMs with journalistic standards and safety considerations.

Paper

References (17)

05Artificial intelligence and journal-ism2019 · Journalism & mass communication quarterly
06Writing for journal-ists2016
072023. Au-tomatingAu-tomating
082023. Towards guidelines for guidelines on the use of generative ai in newsrooms
092023. Navigating the risks of artificial intelligence on the digital news landscape
102023. Data science, machine learning and big data in digital journalism: A survey of state-of-the-art, challenges and opportunitiesExpert Systems with Applications
112023. Gptscore: Evaluate as you desirearXiv preprint
122023. Safety-bench: Evaluating the safety of large language models with multiple choice questionsarXiv preprint

Scroll for more · 5 remaining

Similar papers

© 2026 NYSGPT2525 LLC