Corpora Evaluation and System Bias Detection in Multi-document Summarization

Multi-document summarization (MDS) is the task of reflecting key points from\nany set of documents into a concise text paragraph. In the past, it has been\nused to aggregate news, tweets, product reviews, etc. from various sources.\nOwing to no standard definition of the task, we encounter a plethora of\ndatasets with varying levels of overlap and conflict between participating\ndocuments. There is also no standard regarding what constitutes summary\ninformation in MDS. Adding to the challenge is the fact that new systems report\nresults on a set of chosen datasets, which might not correlate with their\nperformance on the other datasets. In this paper, we study this heterogeneous\ntask with the help of a few widely used MDS corpora and a suite of\nstate-of-the-art models. We make an attempt to quantify the quality of\nsummarization corpus and prescribe a list of points to consider while proposing\na new MDS corpus. Next, we analyze the reason behind the absence of an MDS\nsystem which achieves superior performance across all corpora. We then observe\nthe extent to which system metrics are influenced, and bias is propagated due\nto corpus properties. The scripts to reproduce the experiments in this work are\navailable at https://github.com/LCS2-IIITD/summarization_bias.git.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC