Dear Reviewer,
We sincerely appreciate your thoughtful feedback and welcome the opportunity to address the points you've raised.
- **Concern: The additional training cost seems non-trivial**
Similar to how LLaMA[1] "violates" the scaling law by prioritizing the inference budget, we believe that inference cost is of utmost importance for a context compression method, especially in real-world RAG scenarios involving millions of documents. This principle guided xRAG's design: by reusing existing document embeddings and introducing only a two-layer MLP into the existing LLM, we've minimized additional complexity.
While our method prioritizes inference-time efficiency, it's worth noting that **the training cost is also minimal compared to existing compression methods such as ICAE[2] and LLMLingua[3].** The only trainable component in xRAG is the newly introduced MLP. We would greatly appreciate if you could point us towards any references to context compression methods with trivial training costs that we may have overlooked.
---
- **Concern: The efficiency gain in handling super-long contexts (beyond just 3 documents) is not included.**
RAG is widely recognized as a technique designed to alleviate the need for LLMs to process long contexts by chunking documents into smaller pieces and retrieving the most relevant ones. This is the typical operational model of modern RAG systems. As such, long-context processing and RAG represent two distinct paradigms for LLMs.
Research has shown that useful information generally appears within the top-k documents, and RAG performance tends to plateau as more documents are involved [4] [5]. Consequently, handling super-long contexts falls outside the scope of RAG and, by extension, beyond the scope of xRAG.
However, to address your concern about efficiency with more documents, we've conducted additional tests. If the "super-long" context you mentioned falls within the top-k chunks of RAG (where k is typically less than 10), here are the efficiency results for top-10 documents, following the benchmark setting outlined in Section 5.2 of our paper:
| Top-10 chunks | CUDA Time (s) | GFLOPs | Peak Mem (GiB) |
| ------------- | ------------- | -------------- | -------------- |
| RAG | 3.58s | 10712.22 | 20.42 |
| xRAG | 0.62s (x5.7) | 529.33 (x20.2) | 13.84 (x1.4) |
As demonstrated, xRAG maintains significant efficiency gains even when processing a larger number of documents.
We appreciate your insightful feedback and look forward to further discussion on these points.
---
[1] Touvron, Hugo, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave and Guillaume Lample. “LLaMA: Open and Efficient Foundation Language Models.” ArXiv abs/2302.13971 (2023): n. pag.
[2] Ge, Tao, Jing Hu, Xun Wang, Si-Qing Chen and Furu Wei. “In-context Autoencoder for Context Compression in a Large Language Model.” ArXiv abs/2307.06945 (2023): n. pag.
[3] Jiang, Huiqiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang and Lili Qiu. “LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.” Conference on Empirical Methods in Natural Language Processing (2023).
[4] Lewis, Patrick, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Kuttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel and Douwe Kiela. “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” ArXiv abs/2005.11401 (2020): n. pag.
[5] Shi, Weijia, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Rich James, Mike Lewis, Luke Zettlemoyer and Wen-tau Yih. “REPLUG: Retrieval-Augmented Black-Box Language Models.” ArXiv abs/2301.12652 (2023): n. pag.