Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics

Recent advances in machine learning, particularly in multimodal models, have created new opportunities for analyzing complex data in high-energy physics, where accurate identification of particle interactions is critical for scientific discovery. However, existing approaches rely heavily on convolutional neural networks, which lack interpretability and do not fully leverage multimodal reasoning capabilities. Here we show that a fine-tuned Vision Language Model (VLM) based on LLaMA 3.2 can effectively identify neutrino interactions in pixelated detector data, outperforming both a state-of-the-art convolutional neural network and a Vision Transformer baseline in classification accuracy and robustness. In addition, the VLM provides improved explainability through reasoning-based, interpretable predictions and supports integration of auxiliary semantic information. These results demonstrate the potential of multimodal transformer architectures as general-purpose tools for physics event classification, paving the way for more transparent, flexible, and scalable analysis methods in future high-energy physics experiments. Recent advances in large language models have expanded their capabilities beyond natural language processing, prompting exploration into their applications in high-energy physics. Here, the authors demonstrate that vision language models outperform traditional CNNs in classifying neutrino interactions, offering enhanced accuracy, robustness, and interpretability, thus paving the way for multimodal reasoning in experimental physics.

Paper

Similar papers

© 2026 NYSGPT2525 LLC