Bridging Explainability and Accuracy in CLIP for Object Recognition via Joint Probability Over Rationales and Categories
Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. However, their opaque nature and lack of explainability in predictions make them less trustworthy in critical domains. Providing VLMs with reasonable rationales for object recognition typically comes at the expense of classification accuracy. To tackle this issue, we present ECOR: Explainable CLIP for Object Recognition. First, we propose a mathematical definition of explainability in the object recognition task based on the joint probability distribution of categories and rationales, then leverage this definition to fine-tune CLIP in an explainable manner. Through evaluation on different datasets, ECOR demonstrates state-of-the-art performance in explainable classification. Notably, it excels in zero-shot settings, showcasing its adaptability. This advancement improves explainable object recognition, enhancing trust across diverse applications. We make the code available online (Our code is available on github).
Paper
Full text
Bridging Explainability and Accuracy in CLIP for Object Recognition via Joint Probability Over Rationales and Categories
Semantic Scholar · Computer Science · 2026
Abstract
Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. However, their opaque nature and lack of explainability in predictions make them less trustworthy in critical domains. Providing VLMs with reasonable rationales for object recognition typically comes at the expense of classification accuracy. To tackle this issue, we present ECOR: Explainable CLIP for Object Recognition. First, we propose a mathematical definition of explainability in the object recognition task based on the joint probability distribution of categories and rationales, then leverage this definition to fine-tune CLIP in an explainable manner. Through evaluation on different datasets, ECOR demonstrates state-of-the-art performance in explainable classification. Notably, it excels in zero-shot settings, showcasing its adaptability. This advancement improves explainable object recognition, enhancing trust across diverse applications. We make the code available online (Our code is available on github).
References (65)
Scroll for more · 38 remaining