RedCaps: web-curated image-text data created by the people, for the people

Large datasets of paired images and text have become increasingly popular for\nlearning generic representations for vision and vision-and-language tasks. Such\ndatasets have been built by querying search engines or collecting HTML alt-text\n-- since web data is noisy, they require complex filtering pipelines to\nmaintain quality. We explore alternate data sources to collect high quality\ndata with minimal filtering. We introduce RedCaps -- a large-scale dataset of\n12M image-text pairs collected from Reddit. Images and captions from Reddit\ndepict and describe a wide variety of objects and scenes. We collect data from\na manually curated set of subreddits, which give coarse image labels and allow\nus to steer the dataset composition without labeling individual instances. We\nshow that captioning models trained on RedCaps produce rich and varied captions\npreferred by humans, and learn visual representations that transfer to many\ndownstream tasks.\n

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC