Imagining Grounded Conceptual Representations from Perceptual Information in Situated Guessing Games

In visual guessing games, a Guesser has to identify a target object in a\nscene by asking questions to an Oracle. An effective strategy for the players\nis to learn conceptual representations of objects that are both discriminative\nand expressive enough to ask questions and guess correctly. However, as shown\nby Suglia et al. (2020), existing models fail to learn truly multi-modal\nrepresentations, relying instead on gold category labels for objects in the\nscene both at training and inference time. This provides an unnatural\nperformance advantage when categories at inference time match those at training\ntime, and it causes models to fail in more realistic "zero-shot" scenarios\nwhere out-of-domain object categories are involved. To overcome this issue, we\nintroduce a novel "imagination" module based on Regularized Auto-Encoders, that\nlearns context-aware and category-aware latent embeddings without relying on\ncategory labels at inference time. Our imagination module outperforms\nstate-of-the-art competitors by 8.26% gameplay accuracy in the CompGuessWhat?!\nzero-shot scenario (Suglia et al., 2020), and it improves the Oracle and\nGuesser accuracy by 2.08% and 12.86% in the GuessWhat?! benchmark, when no gold\ncategories are available at inference time. The imagination module also boosts\nreasoning about object properties and attributes.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC