Sports Video Classification Using Vision Transformer: A Deep Learning Based Approach

Sports video classification has become vital for automated sports analytics, highlight generation, and content recommendation. Traditional approaches have relied on handcrafted features or convolutional neural networks (CNN), which have often struggled to capture spatial and temporal dependencies in complex sports footage effectively. To overcome these limitations, we have developed a Vision Transformer (ViT)-based deep learning framework for classifying sports videos into five categories: Cricket, football, basketball, hockey, and tennis. We have constructed a custom sports video dataset that ensures diversity in gameplay scenarios, camera angles, and motion patterns. Our methodology involves pre-processing video frames, extracting spatio-temporal tubelets, and leveraging a transformer-based encoder to learn robust feature representations. Through extensive experiments, our model has achieved an impressive classification accuracy of 98.50 %,demonstrating the superior capability of Vision Transformers to capture complex motion dynamics and distinguish between different sports. The results validate the effectiveness of our approach and highlight its potential for real-world applications in sports analytics, automated content tagging, and intelligent video indexing.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC