Summary
This paper studies the property of the shallow linear Transformer model. The evaluated phenomenon includes comparing Adam and SGD, heavy-tailed gradient noise, condition number, smoothness, etc. The experimental results show that the linear Transformer model reproduces the phenomena that have been observed for full Transformers. The results also show that more heavy-tailed data distribution and more layers can enhance the conclusion.
--------------------------------------------------------
**After rebuttal**: Thank you for your response. I am sorry for the late reply. I am satisfied with the feedback and revisions. I will keep my rating as 6 and support you. A minor point and suggestion: Please clarify what the four figures in Fig. 15 are to be compared within the main body. Later, I realize you are mentioning some figures on page 3. It is not a big issue anyway.
Strengths
This work is novel in terms of new experiments on linear transformers from many aspects of evaluation. The paper is well-written and clear to follow. The studied problem is interesting to the community, from my understanding. The designed experiments are concise to support the conclusion.
Weaknesses
1. This work lacks theoretical understanding or explanation after the experiments in Section 3.2, 4.1, 4.2.
2. The limitation is not discussed. One thing that should be emphasized in the introduction or the abstract is that the experiments are linear regression, which is a good fit for linear Transformers.
Questions
1. One conclusion may not be obvious for linear Transformers compared with softmax Transformers. That is, the attention weights are more concentrated (even sparsely) on some key features/tokens after training, which is observed in several existing works [Li et al., 2023a, Li et al., 2023b, Oymak et al., 2023] for softmax Transformers. Can the authors provide a comparison empirically on this? Also, I think it is better to cover such a discussion in the revision.
Oymak et al., 2023, "On the Role of Attention in Prompt-tuning. "
Li et al., 2023a, "A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity."
Li et al., 2023b, "How do transformers learn topic structure: Towards a mechanistic understanding."
Rating
6: marginally above the acceptance threshold
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.