TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

A crucial component for the scene text based reasoning required for TextVQA\nand TextCaps datasets involve detecting and recognizing text present in the\nimages using an optical character recognition (OCR) system. The current systems\nare crippled by the unavailability of ground truth text annotations for these\ndatasets as well as lack of scene text detection and recognition datasets on\nreal images disallowing the progress in the field of OCR and evaluation of\nscene text based reasoning in isolation from OCR systems. In this work, we\npropose TextOCR, an arbitrary-shaped scene text detection and recognition with\n900k annotated words collected on real images from TextVQA dataset. We show\nthat current state-of-the-art text-recognition (OCR) models fail to perform\nwell on TextOCR and that training on TextOCR helps achieve state-of-the-art\nperformance on multiple other OCR datasets as well. We use a TextOCR trained\nOCR model to create PixelM4C model which can do scene text based reasoning on\nan image in an end-to-end fashion, allowing us to revisit several design\nchoices to achieve new state-of-the-art performance on TextVQA dataset.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC