What Are Tools Anyway? A Survey from the Language Model Perspective

Language models (LMs) are powerful yet mostly for text generation tasks. Tools have substantially enhanced their performance for tasks that require complex skills. However, many works adopt the term "tool" in different ways, raising the question: What is a tool anyway? Subsequently, where and how do tools help LMs? In this survey, we provide a unified definition of tools as external programs used by LMs, and perform a systematic review of LM tooling scenarios and approaches. Grounded on this review, we empirically study the efficiency of various tooling methods by measuring their required compute and performance gains on various benchmarks, and highlight some challenges and potential future research in the field.

Paper

References (88)

Scroll for more · 38 remaining

Similar papers

Reviewer XLAS6/10 · confidence 4/52024-04-20

Summary

This survey paper provides a unified definition of tools as external programs used by language models (LMs) to enhance their abilities. It systematically reviews various scenarios where LMs use tools, including knowledge access, computation, interaction with the real world, and processing non-textual data. The paper also discusses advanced tool usage methods, tool creation approaches, evaluation benchmarks and metrics, and empirically analyzes the trade-offs between performance gains and computation costs when using different tooling methods across tasks.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

- Well-Written and Easy to Understand: The paper is well-organized, clearly written, and easy to follow, even for readers new to the topic of language model tooling. The authors provide clear definitions, illustrative examples, and a logical flow that aids comprehension. - Comprehensive Coverage: The survey comprehensively covers the landscape of language model tooling, including tools for knowledge access, computation, real-world interaction, and multi-modal processing. It examines basic paradigms, advanced methods, benchmarks, and evaluation practices. - Insightful Analysis and Future Directions: Through empirical analysis, the paper identifies efficient tooling approaches and tasks that benefit most from tools. It also critically examines missing evaluation aspects like tool reliability and safety, highlighting important future research directions.

Reasons to reject

My only concern is that, as a survey paper, this work does not present any novel technical contributions or algorithms. Therefore, I am not sure whether it is suitable to publish this paper in COLM?

Reviewer oLaS6/10 · confidence 2/52024-05-05

Summary

This study systematically surveys tools associated with language models (LMs), specifically external programs utilized in LM applications. Beginning with a unified definition of "tools" in the context of LMs, the paper outlines various categories applicable to text-based applications and more sophisticated scenarios. Additionally, the study explores evaluation methods and metrics pertinent to these external programs within the realm of LM.

Rating

6

Confidence

2

Ethics flag

1

Reasons to accept

A unified definition of “tools” is proposed for LMs. A range of relevant works have been listed and categorized in application scenarios. The evaluation of such “tools” has been discussed, particularly the metrics.

Reasons to reject

The contribution of this paper may not be sufficient, considering the coverage and the analysis. For example, retrieval-augmented generation (RAG) can also be viewed as an external tool for LMs, particularly, large language models, to generate more accurate content based on given documents. As for the categorization, it seems to me that “knowledge access”, “Interaction w/ the world”, and “Special-skilled LMs” can be classified as the same functional category, e.g., obtaining external knowledge or information. In addition, readers may be more interested in identifying bottlenecks for such “tools” and how well they can perform on downstream tasks.

Questions to authors

1. Why are the tools in programmatic contexts special and singled out as a section? 2. Can you be more specific on tool creation? Can you give us some guidance on how to create a tool for a specific downstream task based on LMs? 3. What are the benefits of using tools for LMs, compared to other methods like fine-tuning and RAG?

Reviewer m17M8/10 · confidence 5/52024-05-10

Summary

The authors aim to unify various understanding and definitions of "tools" in connection with language models. This aim is highly relevant and needed to guide interested (lay) people as well as researchers and developers and help them to communicate better. The paper contains a nice overview on existing definitions and frameworks and evaluations metrics.

Rating

8

Confidence

5

Ethics flag

1

Reasons to accept

The paper is well written and highly significant. The authors will be able to address the issues raised and submit a revised version.

Reasons to reject

There are no reasons to reject the paper

Questions to authors

What exactly does figure 1 show? The caption does not help very much in understanding it. Please provide a better caption. Additoinally, this figure is not referenced in the text. Please do so and provide contextualizing information to the reader: what are they supposed to see in this figure? Figure 4 is also not referenced in the text. Refering to sections in the paper with the paragraph sign is rather odd. Please use "section 2" etc.

Reviewer H5ew6/10 · confidence 4/52024-05-11

Summary

The paper surveys previous works about LMs with tool-use. Specifically, it proposes a definition of tool-use as calling methods that are executed externally to the LM. It then details about types of tool-use and tool types, and advanced tool use scenarios like tool selection and programmatic contexts. Finally, it presents an empirical investigation of the performance and computation cost of different tool-use approaches.

Rating

6

Confidence

4

Ethics flag

1

Reasons to accept

Tool use is an important subject today when developing language models. As there is a lot of work on the subject, it may be useful to have such surveys in hand. The survey covers different aspects of tool use in a useful high-level perspective.

Reasons to reject

I find the survey and empirical results a bit lacking - I don't think the survey is very comprehensive (e.g. some relevant recent work is missing, for example https://aclanthology.org/2023.findings-emnlp.926/, https://aclanthology.org/2024.eacl-short.10/, https://aclanthology.org/2024.eacl-long.7/). Some earlier works in this area were also missing, e.g. as mentioned in this blog post on the subject: https://newsletter.ruder.io/p/tool-augmented-llms In general, I don't think I learned enough from the survey - it may be somewhat useful for someone completely new to the research area, but I think it can be improved further by discussing more works and diving a bit deeper into how different tool-use methods work and the pros and cons for each.

Questions to authors

typo on page 3 - SPARL section 2.1 - "tools Shumaker et al. (2011)" - missing comma intro - "to obtain time" --> "to obtain the current time"? abstract - "powerful yet mostly" --> "powerful, yet mostly"?

Reviewer m17M2024-06-06

Thanks for the response!

Program Chairsdecision2024-07-10

Decision

Accept

© 2026 NYSGPT2525 LLC