Thank you very much for your patient review. Regarding the two concerns you raised in your response, we provide detailed explanations below.
**Q1: The absence of comparisons with stronger NAR models.**
A1: On the one hand, as far as we know, NAR model is still slightly behind the AR model on most NLG tasks in terms of performance. Therefore, we primarily selected AR for comparison.
On the other hand, certain NAR models like BANG[1] and MIST[2] are pre-trained, making a fair and direct comparison unfeasible. Similarly, some NAR models such as SUNDAE[3] and INSNET[4] have not provided results on datasets like XSUM or IWLST14. The most recent DA-Transformer[5,6] achieved results comparable to AR, but their results were obtained by ensemble the best five checkpoints and using a large beam size of 200. They did not provide results without ensemble, making it difficult for us to compare. Additionally, NAR models like latent-GLAT[7] and CMLMC[8] reported BLEU scores for the IWSLT14 De->En dataset in their papers, respectively, as indicated in the following table.
|Pattern|Model|IWSLT14 De->En|
|:----|:----|:----|
|AR|Transformer|34.74 |
|NAR|GLAT[2021]|29.07|
| |CNAT[2021]|29.81|
| |CMLM[2021]|31.80|
| |latent-GLAT[2022] |32.31|
| |CMLMC[2022]|34.81|
|Diffusion|AR-DIFFUSION ($k$ = 50) |34.95|
| |AR-DIFFUSION ($k$ = 500)|35.62|
As seen from the table above, our method outperforms all the NAR models listed in the table at $k$=50, and its performance is even stronger at $k$=500.
Furthermore, within our paper, Tables 1, 3, and 4 are provided, presenting results across various NAR models (such as CMLM, LevT, CNAT, ConstLeven) for reference.
Nevertheless, we deeply value the importance of your suggestions. We are currently engaged in the pre-training of an AR-Diffusion model. Consequently, in our upcoming version, we intend to implement your recommendations and incorporate comparisons with stronger NAR models.
**Q2: The performance on machine translation benchmarks.**
A2: As addressed in the Q2 of our Rebuttal, the SacreBLEU metric on the translation dataset is indeed slightly lower than AR. However, across other metrics and datasets, we have achieved comparable results with AR. Overall, the performance is on par with AR.
Once again, we truly appreciate your diligent efforts. We hope our response addresses your concerns. Furthermore, we will incorporate all the suggestions you mentioned into the appendix and related work. If you have any further questions, please feel free to reach out to us at your convenience.
[1] Qi W, Gong Y, Jiao J, et al. Bang: Bridging autoregressive and non-autoregressive generation with large scale pretraining[C]//International Conference on Machine Learning. PMLR, 2021: 8630-8639.
[2] Jiang T, Huang S, Zhang Z, et al. Improving non-autoregressive generation with mixup training[J]. arXiv preprint arXiv:2110.11115, 2021.
[3] Savinov N, Chung J, Binkowski M, et al. Step-unrolled Denoising Autoencoders for Text Generation[C]//International Conference on Learning Representations. 2021.
[4] Lu S, Meng T, Peng N. Insnet: An efficient, flexible, and performant insertion-based text generation model[J]. Advances in Neural Information Processing Systems, 2022, 35: 7011-7023.
[5] Huang F, Ke P, Huang M. Directed Acyclic Transformer Pre-training for High-quality Non-autoregressive Text Generation[J]. arXiv preprint arXiv:2304.11791, 2023.
[6] Huang F, Zhou H, Liu Y, et al. Directed acyclic transformer for non-autoregressive machine translation[C]//International Conference on Machine Learning. PMLR, 2022: 9410-9428.
[7] Bao Y, Zhou H, Huang S, et al. latent-GLAT: Glancing at latent variables for parallel text generation[C]//Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022: 8398-8409.
[8] Huang X S, Perez F, Volkovs M. Improving non-autoregressive translation models without distillation[C]//International Conference on Learning Representations. 2021.