This paper explores the properties of back propagation and gradient flow of three basic CNN architectures – VGGNet, ResNet and Inception (GoogLeNet) – in image recognition. For each architecture, we explain the details of forward propagation, the calculation of categorical cross-entropy loss and the backward pass mechanisms. Training stability and transparent gradient monitoring is established in a VGGNet implementation on CIFAR-10. In addition to classical analysis, some modern AI paradigms such as Vision Transformers, self-supervised learning, neural architecture search, and foundation models are revolutionizing the field of image recognition, and are being introduced to augment the capabilities of CNN-based approaches.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex