UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark

Commonsense AI has long been seen as a near impossible goal -- until\nrecently. Now, research interest has sharply increased with an influx of new\nbenchmarks and models.\n We propose two new ways to evaluate commonsense models, emphasizing their\ngenerality on new tasks and building on diverse, recently introduced\nbenchmarks. First, we propose a new multitask benchmark, RAINBOW, to promote\nresearch on commonsense models that generalize well over multiple tasks and\ndatasets. Second, we propose a novel evaluation, the cost equivalent curve,\nthat sheds new insight on how the choice of source datasets, pretrained\nlanguage models, and transfer learning methods impacts performance and data\nefficiency.\n We perform extensive experiments -- over 200 experiments encompassing 4800\nmodels -- and report multiple valuable and sometimes surprising findings, e.g.,\nthat transfer almost always leads to better or equivalent performance if\nfollowing a particular recipe, that QA-based commonsense datasets transfer well\nwith each other, while commonsense knowledge graphs do not, and that perhaps\ncounter-intuitively, larger models benefit more from transfer than smaller\nones.\n Last but not least, we introduce a new universal commonsense reasoning model,\nUNICORN, that establishes new state-of-the-art performance across 8 popular\ncommonsense benchmarks, aNLI (87.3%), CosmosQA (91.8%), HellaSWAG (93.9%), PIQA\n(90.1%), SocialIQa (83.2%), WinoGrande (86.6%), CycIC (94.0%) and CommonsenseQA\n(79.3%).\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC