Response to Reviewer cgEK
We highly appreciate your valuable insights and acknowledgment of our contributions. We hope the following comments could address your concerns.
**(W1): It seems that LTI is only applicable in cases where the edited knowledge can be explicitly represented as knowledge triples $(s, r, o, o^{*})$. However, more complex or nuanced knowledge editing tasks may not fit into this structured format, potentially limiting LTI's applicability in real-world scenarios.**
Thank you for raising this insightful point! We acknowledge that our approach implemented via Multi-stage Constraint, is best suited for edits that can be explicitly represented as knowledge triples (s, r, o, o*), which also the key focus in current research. We agree that enabling models to "Learn To Inference" in more general knowledge contexts is an intriguing topic for future research, and **we've added a dedicated Limitations section in Appendix A of the revised paper to explicitly discuss this limitation.**
**(W2 and Q1): The phenomenon of Editing Overfit and the proposed LTI approach are only tested on small-scale LLMs such as GPT-J and GPT-2. Extending the analysis to larger models, such as LLaMA-2/3 or other larger models, would provide a more comprehensive understanding of LTI’s effectiveness at scale.**
Thank you for your valuable comment! To address the concern, we have extended our experiments to evaluate the performance of baseline methods (FT, FT-L, ROME and MEMIT) and LTI on the **Llama-2-7B** using our EVOKE benchmark (Details results are included in Appendix G), and some results are as follows:
| Editor | Efficacy(↑) | Paraphrase(↑) | Prefix-Distraction DP (↓) | Prefix-Distraction EOS (↑)| Multihop DP (↓) | Multihop EOS (↑) | Subject-Specifity DP (↓) | Subject-Specifity EOS (↑)| Relation-Specificity DP (↓) | Relation-Specificity EOS (↑) |
|-----------------|----------|------------|-------------------|-------------------|------------------|----------------|----------------|-----------------|---------------|----------------|
| Llama-2-7B Base | 13.09 | 15.08 | 16.42 | 58.44 | 1.80 | 86.01 | 0.74 | 97.38 | 1.18 | 82.92 |
| FT | 99.61 | 92.48 | 51.48 | 5.61 | 27.71 | 35.40 | 9.74 | 46.07 | 23.12 | 18.83 |
| FT-L | 92.73 | 21.53 | 16.90 | 53.17 | 2.59 | 85.16 | 1.12 | 95.20 | 1.54 | 81.08 |
| ROME | 100.00 | 93.60 | 22.80 | 46.66 | 21.66 | 63.99 | 21.82 | 45.63 | 2.09 | 81.11 |
| ROME-LTI | 100.00 | 89.42 | 17.90 | 53.13 | 15.55 | 70.32 | 14.30 | 57.42 | 1.73 | 81.36 |
| MEMIT | 100.00 | 96.80 | 40.76 | 25.36 | 30.47 | 48.18 | 39.09 | 23.36 | 3.87 | 78.08 |
| MEMIT-LTI | 100.00 | 91.71 | 26.75 | 36.81 | 14.15 | 67.52 | 9.06 | 47.60 | 2.15 | 80.30 |
From the results, we observe the following conclusions, **which are consistent with our findings on GPT-J and GPT 2-XL models**:
- The baseline methods still demonstrate a pronounced Editing Overfit phenomenon on our EVOKE benchmark. Nearly all successfully edited models exhibit significantly higher Direct Probability (DP) scores compared to the unedited model.
- Both ROME-LTI and MEMIT-LTI demonstrate significant improvements in overfitting metrics (DP and EOS) compared to ROME and MEMIT, indicating effective mitigation of editing overfit.
**(Q2): Are there specific types of facts or domains where the problem of Editing Overfit is more severe? Is there any specific types of knowledge where LTI may be less effective?**
We appreciate the reviewer’s insightful comments! The central point of our paper is that the **Editing Overfit phenomenon likely stems from existing knowledge editing paradigms, which emphasize the direct correspondence between the input prompt $p(s,r)$ and the output $o^*$ for each edit sample $(s,r,o,o^{*})$.** Given that most existing editing methods do not differentiate or focus on knowledge types or domains, and primarily establish input-output mappings, we believe that the Editing Overfit phenomenon likely appears independent of specific knowledge types, while we acknowledge that the extent of overfit could be influenced by knowledge types. We also believe that exploring the performance of various editing models across different types of knowledge domains represents a promising topic for future research.