Exploring Capabilities of Vision Language Models to Perform Explainable Face Image Quality Assessments
The large-scale deployment of face recognition systems in operational environments such as border control, access management, and surveillance makes manual inspection of facial image quality impractical. Automated face image quality assessment (FIQA) algorithms are therefore employed to ensure that only samples of sufficient utility are processed for recognition. While current FIQA methods provide quantitative quality scores or defect-specific measures, these outputs are often difficult to interpret and offer limited actionable guidance to operators or capture subjects. Recent advances in vision-language models (VLMs) have demonstrated strong capabilities in multimodal reasoning and natural language explanation of visual content. In this work, we investigate whether such models can be leveraged to detect face image quality defects and translate them into human-readable, informative feedback. We analyze the ability of state-of-theart VLMs to identify common quality issues and assess their consistency with established FIQA measures.
Paper
Full text
Exploring Capabilities of Vision Language Models to Perform Explainable Face Image Quality Assessments
Semantic Scholar · Computer Science · 2026
Abstract
The large-scale deployment of face recognition systems in operational environments such as border control, access management, and surveillance makes manual inspection of facial image quality impractical. Automated face image quality assessment (FIQA) algorithms are therefore employed to ensure that only samples of sufficient utility are processed for recognition. While current FIQA methods provide quantitative quality scores or defect-specific measures, these outputs are often difficult to interpret and offer limited actionable guidance to operators or capture subjects. Recent advances in vision-language models (VLMs) have demonstrated strong capabilities in multimodal reasoning and natural language explanation of visual content. In this work, we investigate whether such models can be leveraged to detect face image quality defects and translate them into human-readable, informative feedback. We analyze the ability of state-of-theart VLMs to identify common quality issues and assess their consistency with established FIQA measures.