Thank you for your comments, we hope the following helps adding more clarity.
**Sliding Window Length**. Our RULER experiments used 1.4B pre-trained models. The sliding window sizes are: Transformer Baseline — 2048 tokens; B’MOJO, B’MOJO-F and the hybrid baseline — 512 tokens. B’MOJO’s modules never see a sliding window longer than the Transformer’s context length.
**Hybrid baseline**. We have added results in Table B1 in this comment (to complement Table A1 which we copy here). Overall, we find the hybrid model (w/512 token sliding window attention) improves over mamba by a relative 1% @2k and 3% @4k. However, it is slightly weaker than our BMOJO-F (512 token window) and significantly weaker than our full B’MOJO model (512 token window), with a relative performance gap of 7% @2k and 55% @4k.
**Random baseline**. Each task in the RULER benchmark requires generating some specific subset of tokens mentioned in the context, e.g. a 5 digit number. A typical Needle-in-a-Haystack (NIAH) example follows this template: "Some special magic numbers are hidden within the following text. Make sure to memorize it. I will quiz you about the numbers afterwards. \n{context}\n What are all the special magic numbers mentioned in the provided text? The special magic numbers mentioned in the provided text are" (please see Table 2 in the RULER paper). A random baseline (in this case) has to correctly guess 5 numbers out of the vocabulary size, the probability of a correct guess is 1/(10)^5. Harder cases include multiple words, or uuids, which have even lower success probabilities — in practice we measure 0%.
**Concerning Drop w.r.t Transformers**. The drop in performance from using full attention on the 2k context tokens is not concerning, but expected when using a sliding window approach that only leverages 512 tokens. Indeed, as you note above “being able to solve tasks that extend beyond the sliding window length is what is interesting.”; we agree, the Transformer baseline is the paragon in the 2k setting. To further show this, we also evaluate our models on smaller context sizes 512 and 1024, see results in the table below. At size 512, the gap with full attention is indeed null and the gap increases only slightly at size 1024. However, longer contexts set B’MOJO apart from a Transformer model: the latter’s recall performance goes to zero if tested on a context length longer than its attention span, while our models still can recall information from contexts that are up 8x longer than the attention span.
**Extrapolation not often deployed in practice**. Although it is true that often in academic benchmarks, the information supporting the query fits in context, this is not true in many business applications, where the relevant context can be thousands to millions of documents, lines of code, metrics, tables, datasets, and other data that would most definitely not fit in 2048 tokens. With B’MOJO, we are developing a class of models that can cover this long tail of tasks, since the Transformer does not. If in a particular application, 2048 tokens capture the majority of use cases, we would recommend that B’MOJO’s sliding window be set to that value. This way, a practitioner attains the best of both worlds.
**Performance boost**. See [Random baseline] above, and [Concerning Drop w.r.t Transformers]. For Mamba, perhaps looking at relative percentage performance is more revealing. Our B’MOJO model improves over Mamba by 8.5% @2k and 140% @4k relative performance and decreases over transformer by 55% @2k and achieves ~20% accuracy on NIAH at 4k where the transformer model cannot solve the task.
### Table B1: Long context evaluation with RULER (needle in a haystack)
| Context Length | Model | S-NIHA | MK-NIAH | MV-NIAH | MQ-NIAH | Average |
|----------------|--------------|--------|---------|---------|---------|----------|
| 512| Transformer | 100 | 100 | 100 | 100 | 100 |
| | Mamba | 100 | 67 | 78 | 53 | 75 |
| | Hybrid | 100 | 100 | 100 | 100 | 100 |
| | BMOJO-F | 100 | 100 | 100 | 100 | 100 |
| | BMOJO | 100 | 100 | 100 | 100 | 100 |
| 1024| Transformer | 100 | 97 | 63 | 100 | 90 |
| | Mamba | 100 | 44 | 34 | 48 | 57
| | Hybrid| 100 | 53 | 42 | 89 | 71 |
| | BMOJO-F | 100 | 59 | 48 | 98 | 76 |
| | BMOJO | 100 | 81 | 59 | 100 | 85 |
| 2048 | Transformer | 100 | 95 | 62 | 61 | 79 |
| | Mamba | 100 | 32 | 29 | 28 | 47 |
| | Hybrid| 90 | 35 | 34 | 31 | 47.5 |
| | BMOJO-F | 90 | 36 | 35 | 31 | 48 |
| | BMOJO | 90 | 45 | 37 | 33 | 51 |
| 4096 | Transformer | 0 | 0 | 0 | 0 | 0 |
| | Mamba | 9 | 12 | 5 | 7 | 8 |
| | Hybrid | 9 | 13 | 5 | 8 | 8.75 |
| | BMOJO-F | 10 | 16 | 5 | 8 | 10 |
| | BMOJO | 22 | 21 | 17 | 17 | 19 |