<b>Hybrid Transformer Architectures for Robust Sports Classification with Large Language Models</b><b></b>

The rapid increase in sports-related visual data on digital platforms has made classification significantly more difficult, especially among categories with similar scene structures. Traditional CNN-based methods are unable to adequately capture semantic and contextual information in sports classifications that require fine distinctions. This study proposes a new hybrid classification framework that combines Vision Transformer (ViT) architectures, Contrastive Language-Image Pre-training (CLIP), and Explainable Artificial Intelligence (XAI) techniques. The system extracts semantically rich features from ViT-B/16, ViT-B/32, and ViT-L/14 architectures, while examining the impact of these features on classification decisions using the SHAP (SHapley Additive Explanations) method. It then performs the final classification using ensemble machine learning algorithms. The study was tested using two different sports image datasets with 23 and 100 classes. The results show that the ViT-L/14 model achieves an accuracy rate of 99.28% when used with 100 features selected by SHAP. In the 100-class dataset, it achieves an accuracy rate of 98.68%. Throughout all experiments, the proposed method achieved significant improvements in precision, sensitivity, and F1-score metrics while also significantly reducing computation time on low-dimensional feature sets. Additionally, the SHAP-based explainability approach addressed the “black box” issue of deep learning and made the model’s decision-making processes transparent. In conclusion, the proposed ViT+CLIP+XAI-based framework demonstrates superior performance in the classification of large-scale, multi-class sports images and provides fast, accurate, and explainable results. Thanks to its modular structure, it can be easily adapted to has made automated classification increasingly challenging, particularly among categories with similar scene structures. This study proposes a hybrid framework combining Vision Transformer (ViT) architectures, Contrastive Language-Image Pre-training (CLIP), and Explainable Artificial Intelligence (XAI) to address this challenge. Semantically rich features are extracted using ViT-B/16, ViT-B/32, and ViT-L/14 models, ranked via SHAP (SHapley Additive Explanations), and classified through ensemble learning algorithms. Validated on two sports image datasets spanning 23 and 100 classes, the ViT-L/14 model achieves 99.28% accuracy on the 23-class dataset and 98.68% on the 100-class dataset using 100 SHAP-selected features. The proposed method consistently improves precision, recall, and F1-score while significantly reducing computation time on low-dimensional feature sets. Furthermore, SHAP-based explainability addresses the "black box" limitation of deep learning by making model decisions transparent. The proposed ViT+CLIP+XAI framework demonstrates superior performance in large-scale, multi-class sports image classification and is readily adaptable to other complex visual recognition tasks.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC