Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability

Large language models (LLMs) have emerged as powerful tools for many AI problems and exhibit remarkable in-context learning (ICL) capabilities. Compositional ability, solving unseen complex tasks that combine two or more simple tasks, is an essential reasoning ability for Artificial General Intelligence. Despite the tremendous success of LLMs, how they approach composite tasks, especially those not encountered during the pretraining phase, remains an open and largely underexplored question. In this study, we delve into the ICL capabilities of LLMs on composite tasks, with only simple tasks as in-context examples. We develop a test suite of composite tasks including linguistic and logical challenges and perform empirical studies across different LLM families. We observe that models exhibit divergent behaviors: (1) For simpler composite tasks that apply distinct mapping mechanisms to different input segments, the models demonstrate decent compositional ability, while scaling up the model enhances this ability; (2) for more complex composite tasks involving reasoning multiple steps, where each step represents one task, models typically underperform, and scaling up generally provides no improvements. We offer theoretical analysis in a simplified setting, explaining that models exhibit compositional capability when the task handles different input parts separately. We believe our work sheds new light on the capabilities of LLMs in solving composite tasks regarding the nature of the tasks and model scale. Our dataset and code are available at {\url{https://github.com/OliverXUZY/LLM_Compose}}.

Paper

References (77)

Scroll for more · 38 remaining

Similar papers

Reviewer EV4J6/10 · confidence 4/52024-05-06

Summary

This paper is targeting wheter LLMs have compositional ability through in-context learning setting. The paper get intuition from the observation of whether LLMs can solve comsite task formed by simple tasks and make two important contributions to this research question (1) empirically, they introduce several composite tasks to explore how the nature of the seperate task influence compisitonal performance; (2) theoretically, they introduce a simple yet insightful model to provide analysis. Based on these effort, the authors hightlight the variable performance of LLMs on these tasks and theoretical insights into conditions under which LLMs manage to achieve or fail in compositional ability.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

1.This research question, compositional ability, is an important topic in the domain of LLMs and is critical for advancing towards more general AI capabilities. 2.The authors provide clear and insightful empirical experiments of when LLMs succeed or fail on compositional ability across a variety of tasks. 3.The authors have effectively linked the concept of compositional ability with scaling laws, demonstrating that when a model exhibits confined support for simple tasks, enhancements in model scale can yield superior performance. This finding is particularly useful for practical applications.

Reasons to reject

1.The theoretical section lacks a bridge that connects it effectively with the empirical results presented earlier. For instance, the first sentence, 'Despite the complex nature of non-linearity in transformers in LLMs,' does not seem to directly follow from or reference the empirical findings discussed previously 2.Given the extensive and diverse nature of pretraining data (most of them are closed source), there exists a possibility that certain components of the testing tasks may have been indirectly encountered by the models during pretraining phase. So I wonder how do authors ensure this in Section 3: "We make sure composite tasks have not been seen by the LLMs during pretraining.". This should be explained 3.The theoretical conclusions presented in the paper are derived based on certain assumptions regarding the models and data. However, the actual scenarios involving LLMs and their pretraining processes are notably complex. So, my question is whether there is any experimental evidence to validate the application of the proposed theorem for guiding real-world LLMs in tackling combinatorial tasks.

Reviewer sUH46/10 · confidence 4/52024-05-11

Summary

This paper studies the compositional abilities in language models through in-context learning. In this study, the author introduces simple tasks who are quite intuitive to humans to analyze the reasoning abilities of LMs. The authors test these tasks on several LMs like GPT, Llama and Mistral with varying sizes. The results show that (i) models show some compositionality skills which improved with model sizes. (ii) models struggle on tasks requiring sequential reasoning. In addition to empirical analysis the authors also provide theoretical analysis.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

1. The compositional tasks introduced are quite intuitive and can be used to understand compositional abilities of LMs. 2. The experiments are interesting and thorough. Authors also provided reasonable theoretical analysis for the behaviors. 3. The paper is clearly written and easy to follow. I enjoyed reading the paper.

Reasons to reject

None

Questions to authors

1. Given that the tasks are not present in the pretraining data, there is a performance improvement for increased sizes. Do these results correlate to better learning abilities of LMs? 2. Do the authors think these results follow similar trend for all smaller LMs eg: BERT? 3. The authors claim that the tasks are not present in the training data. How do they validate it? 4. The authors missed a couple of related studies in their papers. How does these studies vary from the current proposed one? a. Madasu, Avinash, and Shashank Srivastava. "What do Large Language Models Learn beyond Language?." In Findings of the Association for Computational Linguistics: EMNLP 2022, pp. 6940-6953. 2022. b. Alajrami, Ahmed, Katerina Margatina, and Nikolaos Aletras. "Understanding the Role of Input Token Characters in Language Models: How Does Information Loss Affect Performance?." In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 9085-9108. 2023.

Reviewer Biao6/10 · confidence 3/52024-05-14

Summary

In this paper, the authors investigate the compositional ability of LLMs. They found "(1) For simpler composite tasks that apply distinct mapping mechanisms to different input segments, the models demonstrate decent compositional ability while scaling up the model enhances this ability; (2) for more complex composite tasks that involve reasoning multiple steps, which each step represent one task, models typically underperform, and scaling up does not generally lead to improvements." They offer a theoretical analysis explaining the model when the task handles different input parts separately, the model exhibits compositional capability. This paper provides an angle of how LLM does compositional on a toy task. The paper is written and it is novel. This work may benefit the understanding of LLM.

Rating

6

Confidence

3

Ethics flag

1

Reasons to accept

1. Interesting explanation of how LLM do compositional tasks. 2. Good theoretical analysis.

Reasons to reject

1. Need more evidence on more complex tasks for example mathematical reasonings.

Reviewer YRjW8/10 · confidence 3/52024-05-20

Summary

This paper studies the capabilities of LLMs in solving compositional reasoning tasks -- a crucial capability towards advancement of understanding capabilities of LLMs . They develop a test suite of composite tasks which differ in various linguistic and logical challenges. In the composite logical challenge - they looked at a mix or word and numerical transformations. It was evaluated in 4 settings, the two simple tasks, the composite task with simple tasks as in-context examples and composite task with the composite tasks as in-context example. They find that separable composite tasks, where the simple tasks forming the composite task are distinct, are more straightforward to handle by LLMs, so compose-by-parts tasks seemed easier than compose-by-steps tasks. They also show that composite-in-context examples are easier to handle. In the composite linguistic challenge - they evaluated the ability to translate english based on some customized rules. They constructed two tasks and was also evaluated in the same 4 settings. The authors find that LLMs are capable of handling these composite tasks, and their performance improves with increasing model scale. They attribute this success to the fact that these composite tasks are separable. They support their results through theoretical analysis too where they say for separable composite tasks can occupy different subspaces for the individual simpler tasks in their input embeddings. They also discuss that larger models, with higher rank matrices, are better at compositional capabilities.

Rating

8

Confidence

3

Ethics flag

1

Reasons to accept

This is a well written paper that attempts to understand a very challenging problem for LLMs, i.e. compositional generalization. I like how the authors tried to isolate this task and approached it in a systematic approach. They provide support with both empirical and theoretical findings. Although I still have questions around if the tasks have actually been isolated, I believe the paper's findings could stimulate further research into LLM reasoning and compositional generalization.

Reasons to reject

The approach seems like something that could have been done on top of more commercial and bigger LLMs too, like GPT, LLama-3, Claude and Gemini. I would really like to know how these capabilities scale to those models. Specially they seem to be particularly better at numeric transformations and code-understanding.

Questions to authors

1) I notice that the composite linguistic capabilities generalized much better than the logical ones, could some of the regressions also be attributed to some of the models numerical and code understanding cabilities? 2) When the in-context examples are of the actual composite task instead of the simpler tasks the generalization seems to be much better. I wonder if this implies that the way the in-context examples or the prompt instruction given to the model was tuned. Did the authors try some prompt tuning? 3) If some of the composite logical examples were made more natural language - I wonder if these would have changed things. Specially given my point (1). I'd be curious to know if you tried this?

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC