Preliminary Study of an Evaluation Benchmark for Vision–Language Models in Fashion E-Commerce
We report an evaluation benchmark for assessing the operational suitability of Vision-Language Models (VLMs) in fashion e-commerce. General-purpose benchmarks do not adequately cover fashion-specific attributes or the structured extraction tasks common in e-commerce workflows. We define five tasks across two image streams---outfit and single-item product images---and compare six commercial and two open-source models with multiple prompt variants, including a canonical prompt and model-proposed prompts. Experiments show that the best-performing model varies by task, error patterns are more model-dependent than prompt-dependent, and model updates can improve some tasks while degrading others. These results indicate that task-specific evaluation, prompt robustness checks, and continuous monitoring are practical requirements for deploying VLMs in production fashion systems.
Paper
Full text
Preliminary Study of an Evaluation Benchmark for Vision–Language Models in Fashion E-Commerce
Semantic Scholar · 2026
Abstract
We report an evaluation benchmark for assessing the operational suitability of Vision-Language Models (VLMs) in fashion e-commerce. General-purpose benchmarks do not adequately cover fashion-specific attributes or the structured extraction tasks common in e-commerce workflows. We define five tasks across two image streams---outfit and single-item product images---and compare six commercial and two open-source models with multiple prompt variants, including a canonical prompt and model-proposed prompts. Experiments show that the best-performing model varies by task, error patterns are more model-dependent than prompt-dependent, and model updates can improve some tasks while degrading others. These results indicate that task-specific evaluation, prompt robustness checks, and continuous monitoring are practical requirements for deploying VLMs in production fashion systems.
References (13)
Scroll for more · 1 remaining