An intuitive way to search for images is to use queries composed of an\nexample image and a complementary text. While the first provides rich and\nimplicit context for the search, the latter explicitly calls for new traits, or\nspecifies how some elements of the example image should be changed to retrieve\nthe desired target image. Current approaches typically combine the features of\neach of the two elements of the query into a single representation, which can\nthen be compared to the ones of the potential target images. Our work aims at\nshedding new light on the task by looking at it through the prism of two\nfamiliar and related frameworks: text-to-image and image-to-image retrieval.\nTaking inspiration from them, we exploit the specific relation of each query\nelement with the targeted image and derive light-weight attention mechanisms\nwhich enable to mediate between the two complementary modalities. We validate\nour approach on several retrieval benchmarks, querying with images and their\nassociated free-form text modifiers. Our method obtains state-of-the-art\nresults without resorting to side information, multi-level features, heavy\npre-training nor large architectures as in previous works.\n