Official Rebuttal by Authors
We are grateful for the valuable feedback from all the reviewers. Although some comments appear to be of lower quality and somewhat unfair, some other insights have significantly helped us enhance the paper, taking it to a noticeably higher level in our latest version, Additionally, we also address and clarify some misunderstandings mentioned in the reviews here.
- Comparison with LongWriter (Reviewer UANm, pQCD, srD7): First, LongWriter is a **CONTEMPORARY WORK** with us, released within a time period of two months, so we are not fully eligible to compare with them in all details. Second, we **offer a very clear comparison in the Prerequisite Section Table 1**: 1) We do **not require starting with a robust Long-Context LLM** (we achieve this iteratively), and 2) we can produce responses in a **more diverse form**. Additionally, while LongWriter uses a straightforward agent framework, our approach is superior in terms of **methodological novelty**.
In addition, regarding the evaluation logic similar to LongWriter (Reviewer pQCD), it’s worth noting that a paper contributing a method should not have the eligibility to have a totally different evaluation approach compared to others. For consistent tasks, **evaluations remaining similar is normal**. Moreover, we noticed that the length-following categories in LongWriter are too limited, often just categorized as 'about.' To address this, we've **expanded the categories** to include 'range,' 'about,' 'above,' and 'below,' and we've developed specific scoring metrics for each type. This ensures a more comprehensive length-following evaluation that better aligns with real user needs. We will incorporate this discussion and clearly state our adherence to a similar evaluation logic with LongWriter.
- Explanation on 1/2 and 2/3: We use 1/2 to effectively double the length of the response with each iteration. Our results demonstrate that this approach aligns well with our expectations. The technique of dropping the last part has proven invaluable for seamlessly connecting two parts, and we have found that removing 1/3 is an optimal proportion for achieving this.
- Incorporating GPT-4o for summarizing and analyzing responses to aid human evaluators (Reviewer UANm): Scoring responses entirely by humans is **not practical**. Once you attempt to label an extremely long response, you'll realize the complexity of the task. By utilizing the third-party LLM GPT-4o for summarization and analysis, we can **substantially decrease human effort**. We clarify in the instructions that these summaries and analyses are **for reference only and should not influence their independent judgment.** (Appendix F.1). Additionally, we believe this evaluation method could be adopted in future research on assessing lengthy outputs.
- Adopting same length control for Suri, LongWriter, and ours (Reviewer pQCD): we demonstrate clearly in Section 5.2 that their responses fail to meet appropriate length requirements (Suri) in their queries or apply unsuitable length constraints (LongWriter). If you have doubts, you can verify this by examining their official dataset on HuggingFace. Our evaluation incorporates a much more diverse range of length-following types, rather than focusing solely on terms like 'about.' Additionally, we discover that **implementing the same workflow can considerably enhance their performance**. We **strive to ensure a fair judgment**, and it's unfortunate that some have tried to unjustly accuse us of having ulterior motives.
- Distinct Score Metric (Reviewer UANm, pQCD): The Distinct Score is a well-known metric in NLP used to assess the diversity of responses, introduced by the paper "A Diversity-Promoting Objective Function for Neural Conversation Models."
- Motivation and Methodology Design (Reviewer srD7): See comparisons between Suri and LongWriter.
- Limitations of LLMs in Evaluation (Reviewer srD7): Please check that we also have a human evaluation section.
- Figure 6 display issues (Reviewer srD7): No display issues.
- Lower length control ability (Reviewer pQCD): In the Experiment Section, we **clearly demonstrate that we achieve state-of-the-art performance in length control**. Our results are based on the entire evaluation set. In contrast, the LongWriter paper misleads readers by selectively **testing on a few chosen queries** (LongWrite-Ruler) to produce a seemingly well-aligned figure, which presents an inflated outcome that does not accurately reflect their true capability.
- Acquiring additional details (Reviewer pQCD, UANm, Z4GE): Some of these details have **already been presented** in our paper: We conducted a total of three macro-iterations, as discussed in Sections 5.2 and 6. Additionally, we utilized the initial LLM to validate the generated instructions, discarding any unsuitable ones for producing lengthy responses (Section 4.1). We will clarify these points more clearly and provide more detailed information in our latest version.