ProtoQA: A Question Answering Dataset for Prototypical Common-Sense Reasoning

Given questions regarding some prototypical situation such as Name something\nthat people usually do before they leave the house for work? a human can easily\nanswer them via acquired experiences. There can be multiple right answers for\nsuch questions, with some more common for a situation than others. This paper\nintroduces a new question answering dataset for training and evaluating common\nsense reasoning capabilities of artificial intelligence systems in such\nprototypical situations. The training set is gathered from an existing set of\nquestions played in a long-running international game show FAMILY- FEUD. The\nhidden evaluation set is created by gathering answers for each question from\n100 crowd-workers. We also propose a generative evaluation task where a model\nhas to output a ranked list of answers, ideally covering all prototypical\nanswers for a question. After presenting multiple competitive baseline models,\nwe find that human performance still exceeds model scores on all evaluation\nmetrics with a meaningful gap, supporting the challenging nature of the task.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC