We would like to express our appreciation for your comments. However, we believe there are some misunderstandings and we offer clarifications below.
---
**1. MESA does not affect the three main contributions and findings of this work.**
We would like to emphasize that the three main contributions of our paper remain unaffected by the use of MESA:
- We reveal Mamba’s close relationship to linear attention Transformer.
- We provide detailed analyses of each special design and validate that the forget gate and block design largely lead to Mamba’s superiority.
- We present MLLA, a novel linear attention model that outperforms vision Mamba models.
The first finding is thoroughly analyzed in Section 4 of our paper.
The second one is validated by the results in Table 1 and Table 2, ***where MESA is not employed.***
The last contribution of our work, MLLA, uses MESA in its training. However, as we already mentioned in our previous response, without MESA, MLLA-T can also achieve 83.3 accuracy and still significantly surpasses various vision Mamba models. Notably, the MLLA model is built to validate our last finding, i.e. whether linear attention can match or surpass Mamba in vision, rather than to compete against SOTA vision Transformers.
***In conclusion, MESA does not influence the core contributions of our work, but rather serves as an additional strategy to help MLLA performs optimally.***
We further clarify our intent and the reason for using MESA in the following.
---
**2. The reason for using MESA.**
We would first like to clarify that MESA is just a strategy to prevent overfitting, and it cannot boost performance like token labeling. Token labeling benefits from ***a pre-trained model***, functioning similarly to distillation, whereas MESA only enhances the model's generalization ability and ***doesn't use any pre-trained models.***
Just like in the early stages of studies on visual Transformer, currently vision Mamba research does not have a well-established and universally accepted training protocol. The conventional training setting for vision Transformer may not be optimal for vision Mamba and our Mamba-Like Linear Attention. Therefore, we additionally employ the overfitting prevention strategy MESA to alleviate the overfitting problem of MLLA model and fully demonstrate its potential. Our goal is to provide the community with more robust models.
We believe that excessive pursuit of strictly unchanged training setting could actually restrict the exploration of new architectures.
---
**3. Results without MESA.**
- To better address your request, we further provide the results without MESA.
- Actually, we already provided the result for ***MLLA-T trained without MESA*** in our previous response. Here, we offer a comparison with vision Mamba models based on this result.
| Model | #Params | FLOPs | Acc. |
| :---------------: | :-----: | :---: | :--: |
| Vim-S | 26M | 5.1G | 80.3 |
| VMamba-T | 31M | 4.9G | 82.5 |
| LocalVMamba-T | 26M | 5.7G | 82.7 |
| MLLA-T (w/o MESA) | 25M | 4.2G | 83.3 |
***As you can see, without MESA, MLLA-T can also achieve comparable result and still significantly surpass various SOTA vision Mamba models.***
- We are currently in the process of training ***MLLA-S/B under the no MESA setting*** and will provide those results ***in a few days.***
- Furthermore, we will utilize the backbone trained without MESA to conduct downstream tasks.
- All these results will be included in the revised manuscript. We will provide more experimental results within our capabilities and try our best to benefit the community and follow-up works.
---
**3. Downstream tasks.**
- There seems to be a misunderstanding. In our previous response, we stated that Mask R-CNN 3x results for ***base level model, e.g. VMamba-B,*** is hardly reported by vision Mamba works. However, we did not claim that Mamba-based models have not been tested on 3x schedule. Indeed, we already provided comparison of ***tiny and small level models under 3x schedule in Table 9. of our paper.***
- The primary focus of our paper is on comparisons with vision Mamba models, rather than achieving SOTA results. Given that the works we compared with do not present 3x detection results for base level model, we also omit this experiment previously. Currently, we are working on this experiment to better address your request.