Weaknesses
The main weakness with this paper are overclaiming and lack of citations, which can be misleading for readers. For example, the claim in Figure 1 that "RetNet makes the 'impossible triangle' possible" is an absolute overclaim because the paper lacks validation with larger models and comparison with open-source Transformer models. On the other hand, the authors claim that RWKV and Linear Attention perform poorly, but according to [1], [2], their performance can be on par with Transformers. In Section 2, the authors introduce a new term called "Retention," but this is essentially the same as Linear Attention without the denominator, which has already been proposed in [2], [3]. Additionally, the use of EMA in MEGA and RWKV has been implemented, but the authors fail to cite these works.
[1] Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran G. V., Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartlomiej Koptyra, Hayden Lau, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Xiangru Tang, Bolun Wang, Johan S. Wind, Stanislaw Wozniak, Ruichong Zhang, Zhenyuan Zhang, Qihang Zhao, Peng Zhou, Jian Zhu, and Rui-Jie Zhu. RWKV: reinventing rnns for the transformer era. CoRR, abs/2305.13048.
[2] Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong. The devil in linear transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7025–7041, Abu Dhabi, United Arab Emirates, Dec. 2022. Association for Computational Linguistics.
[3] * Huanru Henry Mao: “Fine-Tuning Pre-trained Transformers into Decaying Fast Weights”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10236–10242, Abu Dhabi, United Arab Emirates, Dec. 2022. Association for Computational Linguistics.
Questions
1. Figure 1 is an absolute overclaim because RWKV[1] and H3[2] have already demonstrated models at the billion-scale level that can achieve performance comparable to Transformers, with parallel training and constant inference. I suggest the authors remove this figure as it could mislead readers.
2. The description of RWKV in Table 1 is completely wrong. According to RWKV[1] and the description in [2], RWKV can indeed be computed in parallel. On the other hand, according to the RWKV paper, its performance is comparable to Transformers, so describing its performance as ✔ is also inaccurate. Overall, Table 1 is highly misleading and can affect the authors' judgment of model performance. I suggest the authors reorganize this table accordingly.
3. The form of Equation 1 is similar to RFA-GATE presented in [7], but the authors did not cite these articles throughout the paper.
4. The form of GroupNorm in Equation 8 is consistent with the NormAttention proposed in [5], but the author does not cite it at all.
5. There is a mistake in the description of the Linear Attention section in Section 2.4. Firstly, Linear Attention refers to the use of the Right-product trick to reduce complexity, but it does not necessarily imply approximation. For example, [4], [5], and [6] do not involve an approximation approach like softmax.
6. The statement "However, linear attention struggles to effectively encode position information, rendering the models less performant" should be supported by a reference since there could be various reasons for the poor performance. Moreover, the issue of the performance of Linear Attention has already been addressed in [5], where they propose a solution.
7. The discussion about MEGA and MEGA-chunk is missing, as they are similar to RetNet in utilizing the EMA technique.
8. Table 2 lacks comparison with open-source models. Firstly, the configuration of the Transformer is not mentioned, whether it is based on GPT2 architecture or Llama architecture. Additionally, there is no information provided regarding the parameter count or training data. On the other hand, there is no comparison with open-source models such as Bloom, Pythia, GPT-Neo, or RWKV. Comparing with these open-source models would allow readers to better understand the performance level of the proposed model.
9. The evaluation scope is too limited, for example, MMLU is not assessed.
10. It is indeed odd that the evaluation datasets in Tables 4 and 5 are inconsistent with Table 3. There should be further explanations provided to clarify this discrepancy.
11. There is a lack of ablation analysis on the architectural design, such as why the dimension of $W_v$ is chosen as d * 2d. On the other hand, an ablation for adding head dimension should be included in Table 5.
[1] Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran G. V., Xuzheng He, Haowen Hou, Przemyslaw Kazienko, Jan Kocon, Jiaming Kong, Bartlomiej Koptyra, Hayden Lau, Krishna Sri Ipsit Mantri, Ferdinand Mom, Atsushi Saito, Xiangru Tang, Bolun Wang, Johan S. Wind, Stanislaw Wozniak, Ruichong Zhang, Zhenyuan Zhang, Qihang Zhao, Peng Zhou, Jian Zhu, and Rui-Jie Zhu. RWKV: reinventing rnns for the transformer era. CoRR, abs/2305.13048.
[2] Tri Dao, Daniel Y. Fu, Khaled Kamal Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré. Hungry hungry hippos: Towards language modeling with state space models. CoRR, abs/2212.14052, 2022.
[3] Eric Martin and Chris Cundy. 2017. Parallelizing linear recurrent neural nets over sequence length. ArXiv, abs/1709.04057.
[4] Zhen Qin, Weixuan Sun, Hui Deng, Dongxu Li, Yunshen Wei, Baohong Lv, Junjie Yan, Lingpeng Kong, and Yiran Zhong. cosformer: Rethinking softmax in attention. In ICLR, 2022.
[5] Zhen Qin, Xiaodong Han, Weixuan Sun, Dongxu Li, Lingpeng Kong, Nick Barnes, and Yiran Zhong. The devil in linear transformer. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7025–7041, Abu Dhabi, United Arab Emirates, Dec. 2022. Association for Computational Linguistics.
[6] Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are rnns: Fast autoregressive transformers with linear attention. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, volume 119 of Proceedings of Machine Learning Research, pages 5156–5165. PMLR, 2020.
[7] Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong. Random feature attention. In International Conference on Learning Representations, 2020