Presentation and Analysis of a Multimodal Dataset for Grounded Language Learning

Grounded language acquisition -- learning how language-based interactions\nrefer to the world around them -- is amajor area of research in robotics, NLP,\nand HCI. In practice the data used for learning consists almost entirely of\ntextual descriptions, which tend to be cleaner, clearer, and more grammatical\nthan actual human interactions. In this work, we present the Grounded Language\nDataset (GoLD), a multimodal dataset of common household objects described by\npeople using either spoken or written language. We analyze the differences and\npresent an experiment showing how the different modalities affect language\nlearning from human in-put. This will enable researchers studying the\nintersection of robotics, NLP, and HCI to better investigate how the multiple\nmodalities of image, text, and speech interact, as well as show differences in\nthe vernacular of these modalities impact results.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC