Toward unsupervised, multi-object discovery in large-scale image collections

This paper addresses the problem of discovering the objects present in a\ncollection of images without any supervision. We build on the optimization\napproach of Vo et al. (CVPR'19) with several key novelties: (1) We propose a\nnovel saliency-based region proposal algorithm that achieves significantly\nhigher overlap with ground-truth objects than other competitive methods. This\nprocedure leverages off-the-shelf CNN features trained on classification tasks\nwithout any bounding box information, but is otherwise unsupervised. (2) We\nexploit the inherent hierarchical structure of proposals as an effective\nregularizer for the approach to object discovery of Vo et al., boosting its\nperformance to significantly improve over the state of the art on several\nstandard benchmarks. (3) We adopt a two-stage strategy to select promising\nproposals using small random sets of images before using the whole image\ncollection to discover the objects it depicts, allowing us to tackle, for the\nfirst time (to the best of our knowledge), the discovery of multiple objects in\neach one of the pictures making up datasets with up to 20,000 images, an over\nfive-fold increase compared to existing methods, and a first step toward true\nlarge-scale unsupervised image interpretation.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC