Predicting Visual Overlap of Images Through Interpretable Non-Metric Box Embeddings

To what extent are two images picturing the same 3D surfaces? Even when this\nis a known scene, the answer typically requires an expensive search across\nscale space, with matching and geometric verification of large sets of local\nfeatures. This expense is further multiplied when a query image is evaluated\nagainst a gallery, e.g. in visual relocalization. While we don't obviate the\nneed for geometric verification, we propose an interpretable image-embedding\nthat cuts the search in scale space to essentially a lookup.\n Our approach measures the asymmetric relation between two images. The model\nthen learns a scene-specific measure of similarity, from training examples with\nknown 3D visible-surface overlaps. The result is that we can quickly identify,\nfor example, which test image is a close-up version of another, and by what\nscale factor. Subsequently, local features need only be detected at that scale.\nWe validate our scene-specific model by showing how this embedding yields\ncompetitive image-matching results, while being simpler, faster, and also\ninterpretable by humans.\n

Paper

References (68)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC