The VISIONE Video Search System: Exploiting Off-the-Shelf Text Search Engines for Large-Scale Video Retrieval
In this paper, we describe in details VISIONE, a video search system that\nallows users to search for videos using textual keywords, occurrence of objects\nand their spatial relationships, occurrence of colors and their spatial\nrelationships, and image similarity. These modalities can be combined together\nto express complex queries and satisfy user needs. The peculiarity of our\napproach is that we encode all the information extracted from the keyframes,\nsuch as visual deep features, tags, color and object locations, using a\nconvenient textual encoding indexed in a single text retrieval engine. This\noffers great flexibility when results corresponding to various parts of the\nquery (visual, text and locations) have to be merged. In addition, we report an\nextensive analysis of the system retrieval performance, using the query logs\ngenerated during the Video Browser Showdown (VBS) 2019 competition. This\nallowed us to fine-tune the system by choosing the optimal parameters and\nstrategies among the ones that we tested.\n