I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing

Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in this field is the establishment of a comprehensive evaluation benchmark for accurately assessing editing results and providing valuable insights for its further development. In response to this need, we propose I2EBench, a comprehensive benchmark designed to automatically evaluate the quality of edited images produced by IIE models from multiple dimensions. I2EBench consists of 2,000+ images for editing, along with 4,000+ corresponding original and diverse instructions. It offers three distinctive characteristics: 1) Comprehensive Evaluation Dimensions: I2EBench comprises 16 evaluation dimensions that cover both high-level and low-level aspects, providing a comprehensive assessment of each IIE model. 2) Human Perception Alignment: To ensure the alignment of our benchmark with human perception, we conducted an extensive user study for each evaluation dimension. 3) Valuable Research Insights: By analyzing the advantages and disadvantages of existing IIE models across the 16 dimensions, we offer valuable research insights to guide future development in the field. We will open-source I2EBench, including all instructions, input images, human annotations, edited images from all evaluated methods, and a simple script for evaluating the results from new IIE models. The code, dataset and generated images from all IIE models are provided in github: https://github.com/cocoshe/I2EBench.

Paper

References (100)

Scroll for more · 38 remaining

Similar papers

Peer review

Reviewer HHca6/10 · confidence 5/52024-06-28

Summary

This paper presents a comprehensive benchmark for instructional image editing. The benchmark contains a high-quality dataset with over 2000 images and 4000 instructions. In addition, the benchmark presents a new evaluation pipeline that leverages GPT to act as the judge to validate the performance of different instructional image editing methods. Within the evaluation, there are two major factor groups considering both high-level and low-level editing.

Strengths

Due to the lack of high-quality benchmarks in image editing, this work fills the gap and is of great significance to the community. The evaluation pipeline is reasonable and clear to follow. The idea of using multimodal large language models to evaluate image editing results is interesting. The overall presentation is clear.

Weaknesses

1. The provided Google Drive link cannot be opened, making the benchmark images inaccessible for review. 2. Image editing evaluation remains a challenge due to few benchmarks. However, the authors ignore comparing their work with a related work [1] due to similar technical pipelines used with that work. Both this and that work first perform human collection, then automated evaluation using GPT, then human evaluation, and then alignment evaluation. The authors should explain the missing comparison. 3. In the high-level editing, the authors did not discuss the evaluation dimension of action change or shape and size change, which are also quite essential editing types. [1] Diffusion Model-Based Image Editing: A Survey. https://arxiv.org/abs/2402.17525

Questions

I am confused about the definition of low-level image editing, which is actually restoration and enhancement. First, low-level restoration tasks usually involve fine-grained operations with rich details. Can GPT really see the slight difference when two similar restored results only vary slightly in terms of visual observation or PSNR evaluation? Second, many editing methods are not designed for low-level vision tasks. It is not appropriate to use them to perform these tasks. The authors should check more methods specially designed for these tasks. I hope the authors can perform additional experiments to address the above concerns. For example, for low-light image enhancement, using two leading methods (such as [2] and [3]) to enhance one image with similar outputs visually (or similar PSNR) and then using GPT for evaluation. [2] Retinexformer: One-stage Retinex-based Transformer for Low-light Image Enhancement (https://arxiv.org/abs/2303.06705) [3] Low-Light Image Enhancement with Wavelet-based Diffusion Models (https://arxiv.org/abs/2306.00306)

Rating

6

Confidence

5

Soundness

2

Presentation

3

Contribution

2

Limitations

The paper should be submitted to the NeurIPS Dataset and Benchmark track.

Authorsrebuttal2024-08-09

Sincere Request for Further Discussions

Dear Reviewer HHca, Thank you for your invaluable efforts and constructive feedback on our manuscript. As the discussion period draws to a close, we eagerly anticipate your thoughts on our response. We sincerely hope that our response meets your expectations. If there are any remaining concerns or aspects that require clarification, we are ready to address them as soon as possible. Best regards, The Authors

Reviewer A2Kh8/10 · confidence 5/52024-06-30

Summary

This paper proposes I2EBench, which is a new benchmark for evaluating Instruction-based Image Editing (IIE) models. It offers a large dataset with over 2,000 images and 4,000 instructions across 16 detailed evaluation dimensions. The benchmark is designed to assess image editing quality automatically and aligns with human perception through extensive user studies. I2EBench aims to provide insights for improving IIE models and will be open-sourced for community use.

Strengths

I commend the authors for their insightful paper, particularly the proposed benchmark, which is remarkably comprehensive and has the potential to significantly advance the field of instruction-based image editing (IIE). 1.The benchmark addresses a wide array of evaluation dimensions, offering a holistic and multifaceted assessment of IIE model capabilities. This comprehensive approach ensures that the strengths and weaknesses of various models are thoroughly examined from multiple perspectives. 2.It strongly emphasizes aligning with human perception, ensuring the benchmark's relevance to human preference. 3.The large collection of images and instructions forms a solid foundation for comprehensive testing. This large and diverse dataset provides a robust foundation for thorough and rigorous testing, enabling models to be evaluated across a broad spectrum of scenarios.

Weaknesses

Although the paper proposes a amazing benchmark, I think the following points can further improve the paper: 1.The results of multiple models on the benchmark were presented in the paper; however, the specifics of their evaluation were not described in sufficient detail. For instance, information regarding the hyperparameters used for the evaluation process and the types of GPUs employed for model testing would provide greater clarity and reproducibility. 2.The rationale behind the use of GPT-4V for supplementary evaluation requires further elucidation. Additionally, whether alternative models could serve as suitable substitutes for GPT-4V? 3.In the section on Instruction, the contributions of the paper are not sufficiently highlighted. A more thorough summarization of the paper's contributions would significantly aid in conveying its impact to the readers.

Questions

I would like to inquire if the I2EBench tool or framework will be open-sourced in the future. This information is crucial for understanding its potential accessibility and contributions to the community.

Rating

8

Confidence

5

Soundness

4

Presentation

4

Contribution

4

Limitations

Yes. The limitations of the paper described by the author in the appendix.

Authorsrebuttal2024-08-09

Sincere Request for Further Discussions

Dear Reviewer A2Kh, Thank you for your invaluable efforts and constructive feedback on our manuscript. As the discussion period draws to a close, we eagerly anticipate your thoughts on our response. We sincerely hope that our response meets your expectations. If there are any remaining concerns or aspects that require clarification, we are ready to address them as soon as possible. Best regards, The Authors

Reviewer A2Kh2024-08-14

Thank you for the author's response. All my concerns have been addressed. After carefully reviewing all the reviewers' comments and the author's response, I found that all reviewers acknowledge the value of I2EBench. I further believe that I2EBench can significantly contribute to the image editing community. Therefore, I will further raise my score and hope the authors can incorporate all the feedback into the revised version.

Authorsrebuttal2024-08-14

Response to Reviewer A2Kh

We greatly appreciate your efforts in reviewing and your recognition of I2EBench's contribution to the community. We are pleased that our response has resolved your concerns.

Reviewer 7DEf8/10 · confidence 5/52024-07-01

Summary

The paper addresses the challenge of evaluating models in the field of Instruction-based Image Editing (IIE) by proposing a comprehensive benchmark called I2EBench. It features: 1) Comprehensive Evaluation: Covers 16 evaluation dimensions for a thorough assessment of IIE models. 2) Human Perception Alignment: Includes extensive user studies to ensure relevance to human preferences. 3) Research Insights: Provides analysis of strengths and weaknesses of existing IIE models to guide future development. I2EBench will be open-sourced, including all instructions, images, annotations, and a script for evaluating new models.

Strengths

1 The supplementary materials provided by the authors offer an in-depth explanation of the evaluation process, which is very beneficial for readers seeking to understand the evaluation details of I2EBench thoroughly. 2 The benchmark proposed in the paper covers a wide range of evaluation dimensions, providing a holistic and multi-faceted assessment of IIE model capabilities. 3 The authors commit to open-sourcing I2EBench, including all relevant resources. This openness will facilitate fair comparisons and knowledge sharing within the community. 4 By systematically evaluating the models, the paper provides valuable research insights that can guide future model architecture design and data selection strategies. 5 The inclusion of user studies in the evaluation process adds depth, ensuring that the evaluation results accurately reflect the real-world experiences of end-users. 6 The benchmark encompasses multiple types of image editing tasks, including both high-level and low-level editing, which makes it versatile and comprehensive. 7 A large number of images and instructions are provided, forming a solid foundation for comprehensive testing of the models.

Weaknesses

1 If the research code or data is not made publicly available, it may limit the usability of the benchmark for other researchers. 2 The paper only provides radar charts for the category experiments. Including the corresponding quantitative data would make the comparisons more precise and understandable. 3 I observed a significant performance gap in the Object Removal dimension between using the original instruction and diverse instruction. The authors could use methods like CLIP similarity or Jaccard similarity to measure the similarity between the original and diverse instructions. This could help determine whether the performance variance is due to significant changes in the instructions or due to the sensitivity of some models to specific vocabulary changes in the Object Removal instructions. 4 The font size in Figure 3 (c) is too small, which might hinder readers from clearly viewing the information presented in the figure.

Questions

1 Why did the authors choose to sample high-level editing images from the COCO dataset, while most low-level editing images are sampled from existing low-level datasets? Others please ref to the weaknesses.

Rating

8

Confidence

5

Soundness

4

Presentation

4

Contribution

4

Limitations

yes

Reviewer riYa5/10 · confidence 5/52024-07-07

Summary

This paper proposes I2EBench, a comprehensive benchmark designed to automatically evaluate the quality of edited images produced by IIE models from multiple dimensions. I2EBench comprises 16 evaluation dimensions, covering both high-level and low-level aspects. Additionally, through user studies, the authors assess the alignment between the proposed benchmark and human perception. The I2EBench dataset consists of over 2000 images for editing, along with corresponding original images and diverse instructions.

Strengths

1. The proposed I2EBench represents a significant advancement over previous works, providing a large and comprehensive benchmark for instruction-based image editing. This contribution will greatly benefit the research community. 2. When establishing I2EBench, the authors have thoughtfully included often overlooked low-level edits, such as rain removal, and have implemented thorough evaluations. 3. The paper is clearly written and easy to follow.

Weaknesses

1. The technical novelty of the evaluation method is insufficient. A significant part of I2EBench's evaluation of high-level edits relies on GPT-4V, such as Direction Perception (line 158) and Object Removal (line 164). Therefore, the authors' contribution appears more like a new prompt engineering method. 2. The insights provided in Section 5 are not informative for the research community. The authors state that "the editing ability across different dimensions is not robust," and Figure 7 shows that current IIE methods perform better on high-level editing tasks (e.g., Object Removal) but struggle with low-level tasks (e.g., Shadow Removal). However, most existing editing datasets are high-level, making it difficult for IIE methods to learn low-level tasks during training. Therefore, the insights in Section 5 seems not constructive.

Questions

1. Does I2EBench consider the aesthetic quality of image edits? For example, in Object Replacement (line 164), how does the benchmark evaluate whether the edited object is appropriately and naturally integrated into the image, rather than simply copied and pasted?

Rating

5

Confidence

5

Soundness

3

Presentation

2

Contribution

3

Limitations

The authors adequately addressed the limitations in the manuscript.

Authorsrebuttal2024-08-09

Sincere Request for Further Discussions

Dear Reviewer riYa, Thank you for your invaluable efforts and constructive feedback on our manuscript. As the discussion period draws to a close, we eagerly anticipate your thoughts on our response. We sincerely hope that our response meets your expectations. If there are any remaining concerns or aspects that require clarification, we are ready to address them as soon as possible. Best regards, The Authors

Reviewer 7DEf2024-08-08

Response to Authors

Thank you for responding to my concerns. For A1, thanks for your commitment. I look forward to the development of I2EBench in the IIE field. For A2, I think quantitative tables can better show the absolute difference in performance than qualitative charts. I suggest replacing Figure 7 with a table. For A3, the response has resolved my issue regarding the performance gap in the object removal dimension. The response is reasonable and supported by experimental evidence. For A4 and A5, thanks for your response. --- This paper proposes a comprehensive benchmark to fill the gap in evaluating high-level and low-level image editing. Based on the author's response, which addressed my concerns, I have decided to increase my score from 6 to 8.

Authorsrebuttal2024-08-09

Response to Reviewer 7DEf

Thank you for acknowledging our work. We will open-source I2EBench in the near future and incorporate your suggestions to further refine our paper.

Reviewer HHca2024-08-10

Thanks for the response. Most of my concerns are addressed. Hence I increase my rating.

Reviewer riYa2024-08-13

Thanks for the response. I lean to keep my rating (5).

Authorsrebuttal2024-08-14

Response to Reviewer riYa

Thank you for your efforts in reviewing and for giving I2EBench a positive rating. Your insightful comments will significantly enhance the paper.

Program Chairsdecision2024-09-25

Decision

Accept (poster)

© 2026 NYSGPT2525 LLC