Experimental results on the Needle-in-a-Haystack task
We appreciate again the acknowledgement for the novelty and insightful contributions of our work. To the best of our knowledge, recent long-context benchmarks like Needle-in-a-Haystack (NIAH), LongBench, and ZeroScrolls typically evaluate instruction-tuned models, as opposed to our pre-trained base models. Nevertheless, we are pleased to share **additional results on the Needle-in-a-Haystack task**.
We found that **Block Transformers perform equally or stronger than loss-equivalent vanilla models**, consistently across **(1) needle locations**, **(2) model scales** and **(3) prompt variants**.
.
## Experimental settings.
Following prior work [1], we construct the context by first sampling 2K-length snippets from concatenated essays written by Paul Graham as the “haystack”, and then inserting a “needle” containing key information in a random location. Following [1], we use this needle format: `The special magic {city} number is: {number}`.
- `{city}` is a randomly chosen city name
- `{number}` is a random 7-digit number.
We then append a prompt that queries to model to retrieve the 7-digit number. We consider two prompt formats:
**1. Gemini prompt**
Format: `<context>\n{context}\n</context>\n\nWhat is the special magic {city} number?\n\nHere is the magic number from the context:`
We mostly followed the NIAH prompt used in Gemini [1], but we excluded the “Don’t give information outside the document or repeat your findings” part, as our models are not instruction-tuned.
**2. Verbatim prompt**
Format: `<context>\n{context}\n</context>\n\n{question}\n\nThe special magic {city} number is:`.
Here, we used the exact same format as that in the needle to query the model.
We measured the accuracy by generating 20 new tokens, and considering a prediction correct if the generated text contains the 7-digit number.
.
## Experimental results.
Note that depth refers to the relative of the location of the needle within the haystack, in percentages.
**Gemini prompt**
| Depth | 0 | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 | Mean |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Vanilla 19M | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.00% | 0.20% | 0.80% | 6.40% | 0.67% |
| Vanilla 85M | 21.00% | 16.40% | 21.80% | 27.40% | 36.60% | 28.00% | 26.80% | 40.20% | 41.80% | 37.60% | 22.80% | 29.13% |
| Vanilla 300M | 46.20% | 69.00% | 72.80% | 78.60% | 76.40% | 70.40% | 71.80% | 74.80% | 73.80% | 78.40% | 66.20% | 70.76% |
| Block 85M | 5.60% | 2.40% | 0.80% | 0.80% | 0.20% | 1.00% | 0.80% | 1.00% | 2.60% | 1.80% | 6.40% | 2.13% |
| Block 300M | 23.40% | 52.60% | 52.60% | 46.60% | 46.00% | 49.20% | 58.40% | 70.40% | 64.00% | 53.60% | 18.40% | 48.65% |
| Block 800M | 35.80% | 74.00% | 76.40% | 78.40% | 69.80% | 77.40% | 76.40% | 79.00% | 75.20% | 72.80% | 53.60% | 69.89% |
| Block 1.2B | 57.20% | 86.60% | 88.80% | 85.60% | 80.40% | 85.20% | 90.40% | 89.20% | 91.00% | 90.40% | 78.80% | 83.96% |
**Verbatim prompt**
| Depth | 0 | 10 | 20 | 30 | 40 | 50 | 60 | 70 | 80 | 90 | 100 | Mean |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| Vanilla 19M | 8.20% | 1.40% | 3.00% | 6.80% | 7.80% | 12.60% | 45.40% | 65.80% | 63.40% | 84.60% | 99.40% | 36.22% |
| Vanilla 85M | 95.60% | 99.40% | 99.00% | 99.40% | 99.20% | 99.20% | 99.00% | 99.60% | 99.60% | 99.00% | 95.60% | 98.60% |
| Vanilla 300M | 99.60% | 100.00% | 100.00% | 99.80% | 100.00% | 100.00% | 99.80% | 100.00% | 100.00% | 100.00% | 99.80% | 99.91% |
| Block 85M | 96.20% | 97.60% | 96.20% | 96.60% | 98.40% | 98.00% | 97.20% | 98.80% | 99.00% | 99.40% | 96.20% | 97.60% |
| Block 300M | 90.20% | 99.40% | 99.60% | 99.20% | 98.60% | 99.60% | 99.60% | 99.80% | 99.80% | 99.20% | 99.20% | 98.56% |
| Block 800M | 95.20% | 99.40% | 98.80% | 98.80% | 98.80% | 98.80% | 99.00% | 99.40% | 99.20% | 97.40% | 99.60% | 98.58% |
| Block 1.2B | 92.60% | 98.40% | 99.40% | 98.80% | 99.60% | 99.60% | 98.80% | 99.80% | 99.80% | 99.20% | 98.00% | 98.55% |
.
These results confirm that the Block Transformer, like the vanilla models, can effectively retrieve global information contained within the 2K context length. With the Gemini prompt, we observed an accuracy trend that was very similar to the perplexity trend of the vanilla vs block models. Near-perfect performance with the Verbatim prompt supports the long-sequence modeling capabilities of our models even when context information is squeeze into a single embedding. We believe this parity between Vanilla and Block Transformers on 2K context length will extend to 8K and beyond.
.
We would appreciate it if you could reflect our additional results on FlashDecoding (modern implementation) and NIAH evaluation (modern evaluation) in your final score, as we believe these have adequately addressed your concerns.
.
[1] Gemini Team, Google. “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.”