We present a neural network model - based on CNNs, RNNs and a novel attention\nmechanism - which achieves 84.2% accuracy on the challenging French Street Name\nSigns (FSNS) dataset, significantly outperforming the previous state of the art\n(Smith'16), which achieved 72.46%. Furthermore, our new method is much simpler\nand more general than the previous approach. To demonstrate the generality of\nour model, we show that it also performs well on an even more challenging\ndataset derived from Google Street View, in which the goal is to extract\nbusiness names from store fronts. Finally, we study the speed/accuracy tradeoff\nthat results from using CNN feature extractors of different depths.\nSurprisingly, we find that deeper is not always better (in terms of accuracy,\nas well as speed). Our resulting model is simple, accurate and fast, allowing\nit to be used at scale on a variety of challenging real-world text extraction\nproblems.\n