FANet: An End-to-End Full Attention Mechanism Model for Multi-Oriented Scene Text Recognition

In this paper, we proposed An end-to-end Multi-Oriented Scene Text Recognition Model with Full Attention Mechanism. Attention mechanism is adopted in both encoder and decoder sides of the model. At the coding end, the idea of residual attention can not only be easily combined with the current most advanced recognition model structure, but also can easily increase the depth of the network without causing model crash, so as to extract image features for encoding more pertinently. The idea of seq2seq attention translation model is adopted in the decoding end, so as to translate image features into recognized words better. This model changes the general method of detecting-slicing-recognition, but directly carries out end-to-end training to get the recognition results. The test results of the model on the two data sets show that the model can achieve even better results using the detecting-slicing-recognition method on the basis of greatly simplifying the model training steps. We call this network FANet.

Paper

Full text

PDF

FANet: An End-to-End Full Attention Mechanism Model for Multi-Oriented Scene Text Recognition

Semantic Scholar · Computer Science · 2019

Abstract

In this paper, we proposed An end-to-end Multi-Oriented Scene Text Recognition Model with Full Attention Mechanism. Attention mechanism is adopted in both encoder and decoder sides of the model. At the coding end, the idea of residual attention can not only be easily combined with the current most advanced recognition model structure, but also can easily increase the depth of the network without causing model crash, so as to extract image features for encoding more pertinently. The idea of seq2seq attention translation model is adopted in the decoding end, so as to translate image features into recognized words better. This model changes the general method of detecting-slicing-recognition, but directly carries out end-to-end training to get the recognition results. The test results of the model on the two data sets show that the model can achieve even better results using the detecting-slicing-recognition method on the basis of greatly simplifying the model training steps. We call this network FANet.

Similar papers

© 2026 NYSGPT2525 LLC