CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward
We propose CrowdVLM-R1, which expands the R1 base model for accurate crowd counting, using a novel framework that integrates the fuzzy group relative policy optimization reward function (FGRPR) to enhance learning efficiency. Unlike the conventional binary (0/1) accuracy reward, the proposed fuzzy reward model, FGRPR, which contains both format and precision rewards, provides nuanced incentives to encourage the $\mathbf{R 1}$ model to learn to adjust policies towards precise outputs. Supervised fine-tuning (SFT) is also integrated for the CrowdVLM-R1 model to learn from a handful of inputs to enable both in-domain and out-of-domain counting. Experimental results demonstrate that GRPO with a standard binary accuracy reward underperforms compared to SFT. In contrast, FGRPR, applied to Qwen2.5-VL-(3B/7B), surpasses all baseline models, including GPT-4o, LLaMA2-70B and SFT, on datasets from different domains. For out-of-domain datasets, FGRPR achieves performance comparable to SFT but excels for tasks with large target values, as its fuzzy reward function assigns higher rewards to approximations closer to the ground-truth. This approach is broadly applicable to tasks where the precision of the answer is critical. The code and data are available at: https://github.com/yeyimilk/CrowdVLM-R1
Paper
References (48)
Scroll for more · 36 remaining