Dilated Convolutional Neural Network for Scenic Text Detection and Recognition on a Synthetic Dataset

Scenic text detection and recognition in natural images pose significant challenges due to complex backgrounds, varying fonts, and irregular layouts. In this research, we propose a novel method utilizing Dilated Convolutional Neural Networks (DCNN) to address these challenges and achieve accurate text detection and recognition on a synthetic dataset The proposed DCNN architecture leverages dilated convolutions to broaden the receptive field of the model while preserving computational efficiency. We generate a large-scale synthetic dataset comprising diverse scenes with synthetic text, annotated with ground-truth bounding boxes and corresponding text labels. We create a sizable synthetic dataset with a variety of sceneries and artificial text that is tagged with real-world bounding boxes and text labels. For text detection, our DCNN model is designed to predict the bounding boxes encompassing text regions within an image. Through extensive experiments, we demonstrate the efficacy of our approach by comparing it with state-of-the-art text detection methods, achieving superior performance in terms of precision, recall, and F1-score. Experimental results on our synthetic dataset and benchmark real-world datasets demonstrate the robustness and generalizability of our proposed method for scenic text detection and recognition. Additionally, qualitative results highlight its capability to handle challenging scenarios, such as curved text, occlusions, and text in various orientations. The proposed model is tested on two standard datasets MNIST and ICDAR03 along with our synthetic dataset. The recognition is evaluated based on the accuracy and our proposed model yields 99.7 percentage of accuracy which is higher when compared to other state of art methods.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC