RULER: What's the Real Context Size of Your Long-Context Language Models?

The needle-in-a-haystack (NIAH) test, which examines the ability to retrieve a piece of information (the"needle") from long distractor texts (the"haystack"), has been widely adopted to evaluate long-context language models (LMs). However, this simple retrieval-based test is indicative of only a superficial form of long-context understanding. To provide a more comprehensive evaluation of long-context LMs, we create a new synthetic benchmark RULER with flexible configurations for customized sequence length and task complexity. RULER expands upon the vanilla NIAH test to encompass variations with diverse types and quantities of needles. Moreover, RULER introduces new task categories multi-hop tracing and aggregation to test behaviors beyond searching from context. We evaluate 17 long-context LMs with 13 representative tasks in RULER. Despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases. While these models all claim context sizes of 32K tokens or greater, only half of them can maintain satisfactory performance at the length of 32K. Our analysis of Yi-34B, which supports context length of 200K, reveals large room for improvement as we increase input length and task complexity. We open source RULER to spur comprehensive evaluation of long-context LMs.

Paper

References (91)

Scroll for more · 38 remaining

Similar papers

Reviewer 4oYN7/10 · confidence 3/52024-04-12

Summary

This paper provides a new synthetic long-context LLM benchmark dubbed RULER. Compared to previous benchmarks, diverse newly designed tasks have been applied with flexible context lengths, characterized by different difficulty levels. Extensive experiments have been conducted on several long-context LLMs under different task settings.

Rating

7

Confidence

3

Ethics flag

1

Reasons to accept

1. A new benchmark with more difficult tasks has been provided. The benchmark could better distinguish the long-context capability of LLMs compared to existing benchmarks, which might contribute to future long-context research. 2. Extensive experiments have been conducted, with deep analysis with demonstration of RULER's capability. 3. Insightful findings on the long-context ability of LLMs might be helpful for future researchers.

Reasons to reject

1. The authors could strictly control variables while synthesizing the dataset. For example, the task of Retrieval, where keys and values are put in the context might be also influential to the performance of LLMs. 2. The synthetic datasets for one task all have the same template, which might not be diverse enough.

Questions to authors

1. How can we make sure the synthetic datasets do not have more preference for certain LLMs, regardless of their long-context learning ability? 2. How to define the complexity of a task? It seems to be more complicated than simply defining it as the length of the context, some other factors might contribute such as distraction level. How to control this factor? 3. For the aggregation task, does LLM on different vocabularies have different performances?

Reviewer 4oYN2024-06-03

Thanks for your considerate response, most of my concerns have been addressed. I would like to see what the additional discussion will provide us in the future version.

Reviewer HSfg7/10 · confidence 4/52024-04-23

Summary

The manuscript introduces RULER, a comprehensive synthetic benchmark designed to evaluate the real-world capabilities of long-context language models (LLMs). The authors argue that traditional benchmarks do not adequately measure true contextual understanding, especially across longer text spans. To address this gap, RULER integrates complex tasks including multi-hop tracing and information aggregation, aiming to provide a more accurate assessment of LLMs' long-context processing abilities.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

1) The evaluation is thorough, testing ten different LLMs across 13 tasks that vary in complexity and length. This comprehensive testing provides a clear picture of how well these models perform under rigorous and varied conditions. 2) The manuscript offers valuable insights into the performance degradation of LLMs as context size increases. These insights are crucial for understanding the limitations of current LLMs and could guide future research and development in the field.

Reasons to reject

1) The dataset constructed in this article primarily extends the range of tasks and enriches the current evaluation dataset. However, it shows limited innovation in terms of novel methodologies or insights into language model capabilities. 2) While RULER's synthetic tasks provide controlled conditions for testing, the absence of a comparative analysis with benchmarks that use realistic data might limit understanding of how these findings translate to real-world applications. 3) It is recommended to supplement the analysis with details on the usage of graphics memory and inference speed for different contexts and models to provide a more comprehensive evaluation of model efficiency.

Reviewer GXoR7/10 · confidence 4/52024-04-30

Summary

This paper presents RULER, a new benchmark to comprehensively evaluate the capabilities of long-context language models through synthetic tasks with flexible configurations. These tasks expand the needle-in-a-haystack test which is currently popular. Empirical results reveal that current long-text models exhibit large degradation on these complex tasks, and training on longer sequences does not always lead to better performance.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

- Providing a novel benchmark that incorporates diverse tasks to test long-context abilities of large language models. - Thorough evaluation and error analysis of a range of models, providing a comprehensive understanding of their performance characteristics. - Well-written paper with clear motivation, illustration of tasks, and presentation of results.

Reasons to reject

- There are many similar benchmarks, and the significance of this paper is not prominent. Further discussion may be needed on the advantages of this paper's benchmark and why these advantages are important. - The synthetic nature of tasks may not fully capture the challenges of real-world long-context applications. It may sound easy to construct more complex synthesis tasks, but we need more insights about real long-context requirements. For example, How well do the synthetic tasks in RULER correlate with performance on real-world long-context applications? Is it necessary for long-context models to complete these complex tasks such as extreme long chain of binding statements?

Reviewer cgKY7/10 · confidence 4/52024-05-10

Summary

The paper introduces a new synthetic benchmark, RULER, designed to evaluate the long-context modeling capabilities of language models (LMs). The authors argue that existing tests, such as the needle-in-a-haystack (NIAH), are insufficient to comprehensively assess LMs' understanding of long contexts. RULER expands beyond simple retrieval tasks to include multi-hop tracing, aggregation, and question answering, allowing for a more nuanced evaluation of model performance across various context lengths and complexities. The study involves testing ten long-context LMs with 13 representative tasks and finds that despite near-perfect accuracy in NIAH, all models experience significant performance drops as context length increases. The paper's analysis reveals that only a few models can maintain satisfactory performance at 32K tokens, and there is substantial room for improvement, especially as input length and task complexity increase.

Rating

7

Confidence

4

Ethics flag

1

Reasons to accept

1. RULER provides a more thorough evaluation of LMs by including diverse tasks that go beyond simple retrieval, offering a better understanding of how models handle long contexts. The benchmark allows for adjustments in sequence length and task complexity, making it adaptable for various research needs and future model evaluations. 2. The paper presents detailed results and analysis across different models and context sizes, giving insights into the limitations and performance degradation of current LMs. 3. The study reveals that most models struggle with complex tasks in long contexts, despite their claims of handling large context sizes, which is valuable information for both researchers and practitioners.

Reasons to reject

1. The study focuses on a selection of models, which might not represent the entire spectrum of available LMs, possibly missing insights from a broader set. Actually, there are lots of models that specially optimized for long context scenarios, such as Claude and Moonshot, which the performance needs to be discussed.

Questions to authors

na

Reviewer GXoR2024-06-05

Thank you very much for your responses. I will keep my current rating score and tend to accept the paper.

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC