CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language Learning

Approaches to Grounded Language Learning typically focus on a single\ntask-based final performance measure that may not depend on desirable\nproperties of the learned hidden representations, such as their ability to\npredict salient attributes or to generalise to unseen situations. To remedy\nthis, we present GROLLA, an evaluation framework for Grounded Language Learning\nwith Attributes with three sub-tasks: 1) Goal-oriented evaluation; 2) Object\nattribute prediction evaluation; and 3) Zero-shot evaluation. We also propose a\nnew dataset CompGuessWhat?! as an instance of this framework for evaluating the\nquality of learned neural representations, in particular concerning attribute\ngrounding. To this end, we extend the original GuessWhat?! dataset by including\na semantic layer on top of the perceptual one. Specifically, we enrich the\nVisualGenome scene graphs associated with the GuessWhat?! images with abstract\nand situated attributes. By using diagnostic classifiers, we show that current\nmodels learn representations that are not expressive enough to encode object\nattributes (average F1 of 44.27). In addition, they do not learn strategies nor\nrepresentations that are robust enough to perform well when novel scenes or\nobjects are involved in gameplay (zero-shot best accuracy 50.06%).\n

Paper

References (55)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC