Meta-review
This paper proposes the QSD-Transformer which is a quantized framework for spike-based Transformers by introducing a bi-level optimization strategy, incorporating information-enhanced LIF, and fine-grained distillation to rectify the attention distribution. Most of the reviewers have positive comments on this work and part of them raised their score after the response. Thus, this paper can be accepted and please prepare the final version well.
Additional comments on reviewer discussion
Reviewer xmWz (Point Raised):
Concern over the potential reduction in training and inference efficiency due to the extension to 4 virtual timesteps.
Authors’ Response:
Clarified that the use of multi-bit pulses within a single time step during training (IE-LIF) and their conversion to binary pulses during inference reduces memory requirements and improves training speed. Provided data showing a 3.2× speedup in training and a 6.1× reduction in memory usage.
Reviewer xmWz (Point Raised):
Suggestion for a more thorough empirical comparison with other state-of-the-art methods.
Authors’ Response:
Included results comparing QSD-Transformer with the latest state-of-the-art method, QKformer, demonstrating the effectiveness of their approach.
Reviewer pkTr (Point Raised):
Lack of comparison with previous SNN and quantized ANN transformer models.
Authors’ Response:
Added comparisons with QKFormer and QAT-ViT models to strengthen the empirical evidence.
Reviewer pkTr (Point Raised):
Concern about the scalability of the method, given the use of small architectures for ImageNet experiments.
Authors’ Response:
Explained that the small architectures were the result of quantizing larger spike-based Transformer models and that the largest quantization baseline used was 55M, demonstrating scalability. Plans to apply the method to even larger models in the future were mentioned.
Reviewer pkTr (Point Raised):
Large training overhead due to multi-bit spikes and knowledge distillation.
Authors’ Response:
Provided analysis showing that IE-LIF and FGD techniques did not increase training overhead and in fact reduced it compared to traditional spike-based Transformers.
Reviewer pkTr (Point Raised):
Insufficient information in the transfer learning section, specifically about bit-width and accuracy without transfer learning.
Authors’ Response:
Clarified the bit-width used and provided accuracy results for CIFAR10/100 and CIFAR10-DVS with direct training, showing high accuracy.
Reviewer pkTr (Point Raised):
Request for firing rate information and changes in the self-attention part compared to the original Spike-driven Transformer-V2.
Authors’ Response:
Apologized for the oversight and presented the changes in the firing rates of the attention mechanism modules before and after quantization.
Reviewer xmWz:
Request for comparative experimental results with Fast-SNN.
Authors’ Response:
Agreed to add comparative experimental results with Fast-SNN, explaining how Fast-SNN inspired their IE-LIF approach, which uses multi-bit values during training and binary spikes during inference.
Reviewer xmWz :
Inquiry about the training efficiency compared to the original spiking transformer.
Authors’ Response:
Confirmed that the QSD-Transformer quantization approach improves training efficiency due to the use of IE-LIF neurons, which reduce training time and memory consumption.
Reviewer xmWz:
Feasibility of applying the method to NLP tasks.
Authors’ Response:
Confirmed the feasibility and added experiments on NLP tasks using the QSD-Transformer quantization framework, with comparisons based on Spikezip and SpikeBERT.
Reviewer pkTr :
Concern that IE-LIF neurons do not utilize temporal information, unsuitable for temporal benchmarks.
Authors’ Response:
Clarified that IE-LIF enables multi-time-step forward propagation during training, suitable for temporal benchmarks, and provided evidence of improved performance on CIFAR10-DVS.
Reviewer pkTr :
Limited comparison between ANN2SNN and Direct Training methods, with other methods outperforming QSD-Transformer.
Authors’ Response:
Addressed the concern by including a comparison with state-of-the-art ANN2SNN methods like SpikeZip and ECMT, and demonstrated the application of QSD-Transformer to models like Spikezip, showing improved results.
The authors effectively addressed the concerns and questions raised by the reviewers, providing additional experimental results and clarifications that strengthened the manuscript.
The decision to accept the paper was influenced by the authors’ ability to show improved training efficiency, the feasibility of applying their method to NLP tasks, and the suitability of their approach for temporal benchmarks.
The inclusion of comparative results with Fast-SNN and other state-of-the-art methods demonstrated the robustness of the QSD-Transformer quantization framework.
The decision also considered the overall contribution of the work to the field and the clarity of the responses to the reviewers’ points.