This paper presents a new model for the task of scene text visual question\nanswering, in which questions about a given image can only be answered by\nreading and understanding scene text that is present in it. The proposed model\nis based on an attention mechanism that attends to multi-modal features\nconditioned to the question, allowing it to reason jointly about the textual\nand visual modalities in the scene. The output weights of this attention module\nover the grid of multi-modal spatial features are interpreted as the\nprobability that a certain spatial location of the image contains the answer\ntext the to the given question. Our experiments demonstrate competitive\nperformance in two standard datasets. Furthermore, this paper provides a novel\nanalysis of the ST-VQA dataset based on a human performance study.\n