The transformer architecture, originally used in the NLP tasks, is recently applied in computer vision tasks. The accuracy is improved when taking the ViT block into consideration. However, the improvement is limited since the Vit focuses on local details and breaks the global coherence of the image. In this paper, progressive multi-step training and knowledge distillation strategies are adopted in the vision transformer, preserving and providing more information from the macro view. The model progressively learns the information stored in the image crossing multiple granularities by adopting progressive training. After processing the whole image, the learned identifiable features flow into the self-distillation framework making it possible for massive data scale tasks. The experiments are conducted on several generally used datasets, and the output performance achieves a distinctive increase in accuracy when working on the CUB, CAR and AIR datasets compared to current state-of-the-art methods.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex