Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

While Ferret seamlessly integrates regional understanding into the Large Language Model (LLM) to facilitate its referring and grounding capability, it poses certain limitations: constrained by the pre-trained fixed visual encoder and failed to perform well on broader tasks. In this work, we unveil Ferret-v2, a significant upgrade to Ferret, with three key designs. (1) Any resolution grounding and referring: A flexible approach that effortlessly handles higher image resolution, improving the model's ability to process and understand images in greater detail. (2) Multi-granularity visual encoding: By integrating the additional DINOv2 encoder, the model learns better and diverse underlying contexts for global and fine-grained visual information. (3) A three-stage training paradigm: Besides image-caption alignment, an additional stage is proposed for high-resolution dense alignment before the final instruction tuning. Experiments show that Ferret-v2 provides substantial improvements over Ferret and other state-of-the-art methods, thanks to its high-resolution scaling and fine-grained visual processing.

Paper

References (71)

Scroll for more · 38 remaining

Similar papers

Reviewer aBcC6/10 · confidence 4/52024-05-10

Summary

The paper introduces Ferret-v2, an enhancement of the Ferret model, aimed at refining the capabilities of multimodal large language models (MLLMs) for high-resolution image understanding. The authors focus on improving the model's referring and grounding abilities, particularly in detailed regional and global reasoning. The introduction of a multi-granularity visual encoding and a three-stage training paradigm are novel and effective. The results, compared against other state-of-the-art methods, demonstrate substantial improvements.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

(1) The paper is well-written, and clear, providing extensive empirical evidence to support its claims. (2) The approach integrates innovative techniques like any resolution grounding, multi-granularity visual encoding, and a novel training paradigm, which are well justified and rigorously tested. (3) The paper provides a comprehensive set of experiments and comparisons to benchmark datasets, demonstrating the effectiveness and efficiency of Ferret-v2.

Reasons to reject

(1) The paper briefly mentions potential risks of harmful and counterfactual responses from MLLMs but does not thoroughly explore limitations or potential biases of the proposed model. (2) The complexity of the model and the training process might pose challenges for reproducibility, which are not adequately addressed in the manuscript.

Questions to authors

(1) Could you elaborate on the potential limitations and ethical considerations associated with Ferret-v2, especially concerning its deployment in real-world applications? (2) How much does the model training costs? Will the code and model publicly released?

Reviewer 889e6/10 · confidence 4/52024-05-10

Summary

This paper is an improved version of Ferret. The improvement includes high image resolution, multi-granularity visual encoding and three-stage training pipeline. These techniques bring a lot benefit to Ferret and result to leading performance in various multimodal benchmarks, especially on region-level tasks.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

This paper introduces a series of effective techniques to improve the multimodal model, - (1) high image resolution: sub-patches and direct upsampling are well studied. - (2) multi-granularity visual encoding: taking advantages of CLIP and DINOv2 on global image feature extraction and low-level details preservation, respectively. - (3) multi-stage training pipeline: I really appreciate that Stage III unfreezes CLIP and DINOv2, which is a very brave attempt and rarely seen in recent multimodal works. By integrating these to Ferret, leading experimental results are produced in both image-level and region-level multimodal benchmarks.

Reasons to reject

The main concern is that this is an incremental paper from the perspective of technical originality. For example, the sub patches-based high resolution has been used in [1][2], the integration of CLIP and DINO as the vision encoder are studied in [3], the multi-stage training strategies are widely used in previous multimodal works. This is the first COLM, the acceptance threshold is not clear to me. If technical originality is a highly strict requirement of COLM, Ferret-v2 may not be novel enough to be accepted. However, if a well-constructed engineering paper is recommended, Ferret-v2 is truly a good example. My initial score is a conservative score. I will determine my final score after reading other reviewers comments. Reference: - [1] LLaVA-NeXT: Improved reasoning, OCR, and world knowledge - [2] Towards Open-Ended Visual Recognition with Large Language Model - [3] From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reviewer NGY17/10 · confidence 4/52024-05-11

Summary

The authors propose Ferret-v2, an improved baseline of Ferret. Ferret-v2 has three key designs: any resolution processing, multi-granularity visual encoding, and a three-stage training paradigm. The referring and grounding abilities of Ferret-v2 are strengthened by these designs. Finally, experiments demonstrated the effectiveness of the proposed Ferret-v2 on referring, grounding tasks, and modern MLLM benchmarks.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

1. The authors propose an improved Ferret-v2 with three specialized designs to strengthen the referring and grounding capabilities. Ferret-v2 seems novel to me. 2. The proposed Ferret-v2 is effective and superior on multiple benchmarks, including referring tasks, grounding tasks, and modern MLLM benchmarks. 3. This paper is easy to follow.

Reasons to reject

1. Efficiency and scalability. How does Ferret-v2 balance accuracy with computational efficiency? The time complexity of Ferret-v2 is suggested to be clarified. 2. Potential harmful outputs and ethical considerations. The authors are suggested to discuss how Ferret-v2 prevents the potentially harmful outputs.

Questions to authors

please see the ''reasons to reject''

Reviewer NGY12024-06-04

Response to rebuttal

I would like to thank the authors's response and stick to keeping my rating.

Authorsrebuttal2024-06-06

Thank you for the feedbacks

Dear Reviewer NGY1, Thank you for your valuable feedback, and we really appreciate the effort and time you took to review our submission. We will address the above concerns in our paper revision. Warm regards, Author(s)

Authorsrebuttal2024-06-06

Reviewer-Author Discussion Period Ends in ONE Day

Dear Reviewer 889e, Thank you again for your insightful reviews of our submission. Following your feedback, we have provided a detailed response trying to address the concerns you raised. As the deadline is approaching, it would be very helpful if you could revisit our clarifications and let us know if any ambiguities remain before the reviewer-author discussion period ends. We would greatly appreciate any further comments you might have regarding our work, and we are fully committed to answering any questions. Your effort and time in reviewing our submission are sincerely appreciated. Warm regards, Author(s)

Reviewer 889e2024-06-07

I really appreciate the authors' response. I understand there are indeed technical differences between Ferret-v2 and previous works, but these differences are still incremental compared with Ferret-v1. After reading other reviewers' opinions and scores, I keep my score. Good luck !

Authorsrebuttal2024-06-06

Reviewer-Author Discussion Period Ends in ONE Day

Dear Reviewer aBcC, Thank you again for your insightful reviews of our submission. Following your feedback, we have provided a detailed response trying to address the concerns you raised. As the deadline is approaching, it would be very helpful if you could revisit our clarifications and let us know if any ambiguities remain before the reviewer-author discussion period ends. We would greatly appreciate any further comments you might have regarding our work, and we are fully committed to answering any questions. Your effort and time in reviewing our submission are sincerely appreciated. Warm regards, Author(s)

Reviewer aBcC2024-06-07

Thank you for your reply. I have carefully read the response and I will keep my original ratings.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC