Visual grounding tasks aim to localize image regions based on natural language references. In this work, we ex-plore whether generative VLMs predominantly trained on image-text data could be leveraged to scale up the text an-notation of visual grounding data. We find that grounding knowledge already exists in generative VLM and can be elicited by proper prompting. We thus prompt a VLM to generate object-level descriptions by feeding it object regions from existing object detection datasets. We fur-ther propose attribute modeling to explicitly capture the im-portant object attributes, and spatial relation modeling to capture inter-object relationship, both of which are common linguistic pattern in referring expression. Our constructed dataset (500K images, 1M objects, 16M referring expressions) is one of the largest grounding datasets to date, and the first grounding dataset with purely model-generated queries and human-annotated objects. To verify the qual-ity of this data, we conduct zero-shot transfer experiments to the popular RefCoco benchmarks for both referring expression comprehension (REC) and segmentation (RES) tasks. On both tasks, our model significantly outperform the state-of-the-art approaches without using human anno-tated visual grounding data. Our results demonstrate the promise of generative VLM to scale up visual grounding in the real world.
Paper
References (74)
Scroll for more · 38 remaining