Video classification has wide applications in surveillance, sports analysis, content recommendation, and more, making it is important to model the spatial and temporal dynamics accurately. The standard convolutional neural networks (CNNs) have proven to be excellent in spatial feature extraction but are less competent in modeling temporal relationships. This is one of the biggest problems with these models not being able to model long-range dependencies, which has led to the development of Vision Transformers (ViTs). However, their application in video classification remains underexplored. In this paper, the problem of improving the accuracy of video classification is approached by combining the strengths of ViTs and traditional classifiers. A pre-trained Vision Transformer is used to obtain high-dimensional feature representations, which are then integrated with a Support Vector Machine (SVM) that performs the classification. But the results of our experiment show that our ViT-SVM combination actually overcomes traditional methods and is thus a fairly competitive approach with a higher accuracy rate. These results suggest the possibility of using hybrid models to improve the accuracy of video classification.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex