Bridging Explainability and Accuracy in CLIP for Object Recognition via Joint Probability Over Rationales and Categories

Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. However, their opaque nature and lack of explainability in predictions make them less trustworthy in critical domains. Providing VLMs with reasonable rationales for object recognition typically comes at the expense of classification accuracy. To tackle this issue, we present ECOR: Explainable CLIP for Object Recognition. First, we propose a mathematical definition of explainability in the object recognition task based on the joint probability distribution of categories and rationales, then leverage this definition to fine-tune CLIP in an explainable manner. Through evaluation on different datasets, ECOR demonstrates state-of-the-art performance in explainable classification. Notably, it excels in zero-shot settings, showcasing its adaptability. This advancement improves explainable object recognition, enhancing trust across diverse applications. We make the code available online (Our code is available on github).

Paper

Full text

PDF

Bridging Explainability and Accuracy in CLIP for Object Recognition via Joint Probability Over Rationales and Categories

Semantic Scholar · Computer Science · 2026

Abstract

Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. However, their opaque nature and lack of explainability in predictions make them less trustworthy in critical domains. Providing VLMs with reasonable rationales for object recognition typically comes at the expense of classification accuracy. To tackle this issue, we present ECOR: Explainable CLIP for Object Recognition. First, we propose a mathematical definition of explainability in the object recognition task based on the joint probability distribution of categories and rationales, then leverage this definition to fine-tune CLIP in an explainable manner. Through evaluation on different datasets, ECOR demonstrates state-of-the-art performance in explainable classification. Notably, it excels in zero-shot settings, showcasing its adaptability. This advancement improves explainable object recognition, enhancing trust across diverse applications. We make the code available online (Our code is available on github).

References (65)

Scroll for more · 38 remaining

Similar papers

© 2026 NYSGPT2525 LLC