Player Image Pre-trained Vision Transformer for Rugby Scene Classification

Scene classification is a crucial task in sports analytics, enabling automated recognition of key game events for tactical analysis and highlight generation. Traditional approaches often struggle to capture fine-grained player-state information essential for accurate scene interpretation. In this work, we propose a self-supervised learning approach that leverages player-specific image crops extracted from sports videos to enhance scene classification performance. We apply object detection to isolate athlete regions and construct a domain-specific dataset for pretraining a Vision Transformer using a masked autoencoder framework. The resulting PlayerImage-ViT model learns player-centric representations that significantly improve scene classification accuracy. Evaluation on tackle scene classification demonstrates a 41.95% improvement over ImageNet-pretrained baseline models. To assess feature efficiency, we conduct a comprehensive dimensionality analysis using principal component analysis on the 768-dimensional PlayerImage-ViT embeddings. Our analysis reveals that while a single principal component provides substantial performance gains, optimal results are achieved with 20-100 components depending on the backbone architecture, with ResNet-50 showing peak performance at 100 dimensions. Visualization of learned features confirms that our model effectively highlights player regions critical for human scene perception. These results demonstrate the effectiveness of athlete-centric representation learning for automated sports analytics and provide insights into optimal feature dimensionality for practical deployment.

Paper

Full text

PDF

Player Image Pre-trained Vision Transformer for Rugby Scene Classification

OpenAlex · Video Analysis and Summarization · 2025

Abstract

Scene classification is a crucial task in sports analytics, enabling automated recognition of key game events for tactical analysis and highlight generation. Traditional approaches often struggle to capture fine-grained player-state information essential for accurate scene interpretation. In this work, we propose a self-supervised learning approach that leverages player-specific image crops extracted from sports videos to enhance scene classification performance. We apply object detection to isolate athlete regions and construct a domain-specific dataset for pretraining a Vision Transformer using a masked autoencoder framework. The resulting PlayerImage-ViT model learns player-centric representations that significantly improve scene classification accuracy. Evaluation on tackle scene classification demonstrates a 41.95% improvement over ImageNet-pretrained baseline models. To assess feature efficiency, we conduct a comprehensive dimensionality analysis using principal component analysis on the 768-dimensional PlayerImage-ViT embeddings. Our analysis reveals that while a single principal component provides substantial performance gains, optimal results are achieved with 20-100 components depending on the backbone architecture, with ResNet-50 showing peak performance at 100 dimensions. Visualization of learned features confirms that our model effectively highlights player regions critical for human scene perception. These results demonstrate the effectiveness of athlete-centric representation learning for automated sports analytics and provide insights into optimal feature dimensionality for practical deployment.

References (61)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC