IMDB-WIKI-SbS: An Evaluation Dataset for Crowdsourced Pairwise\n Comparisons

Today, comprehensive evaluation of large-scale machine learning models is\npossible thanks to the open datasets produced using crowdsourcing, such as\nSQuAD, MS COCO, ImageNet, SuperGLUE, etc. These datasets capture objective\nresponses, assuming the single correct answer, which does not allow to capture\nthe subjective human perception. In turn, pairwise comparison tasks, in which\none has to choose between only two options, allow taking peoples' preferences\ninto account for very challenging artificial intelligence tasks, such as\ninformation retrieval and recommender system evaluation. Unfortunately, the\navailable datasets are either small or proprietary, slowing down progress in\ngathering better feedback from human users. In this paper, we present\nIMDB-WIKI-SbS, a new large-scale dataset for evaluating pairwise comparisons.\nIt contains 9,150 images appearing in 250,249 pairs annotated on a\ncrowdsourcing platform. Our dataset has balanced distributions of age and\ngender using the well-known IMDB-WIKI dataset as ground truth. We describe how\nour dataset is built and then compare several baseline methods, indicating its\nsuitability for model evaluation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC