Large datasets of paired images and text have become increasingly popular for\nlearning generic representations for vision and vision-and-language tasks. Such\ndatasets have been built by querying search engines or collecting HTML alt-text\n-- since web data is noisy, they require complex filtering pipelines to\nmaintain quality. We explore alternate data sources to collect high quality\ndata with minimal filtering. We introduce RedCaps -- a large-scale dataset of\n12M image-text pairs collected from Reddit. Images and captions from Reddit\ndepict and describe a wide variety of objects and scenes. We collect data from\na manually curated set of subreddits, which give coarse image labels and allow\nus to steer the dataset composition without labeling individual instances. We\nshow that captioning models trained on RedCaps produce rich and varied captions\npreferred by humans, and learn visual representations that transfer to many\ndownstream tasks.\n
Paper
References (100)
Scroll for more · 38 remaining