When Do Tools and Planning Help Large Language Models Think? A Cost- and Latency-Aware Benchmark
Modern large language models (LLMs) increasingly rely on inference-time planning and external tools to improve reasoning. We benchmark this behavior on two real-world settings: event-centric question answering over graph-structured knowledge (Event-QA) and persuasive response generation in Reddit ChangeMyView (CMV). Using LangChain and LangGraph, we compare a one-shot baseline against a plan-execute-replan agent equipped with task-specific tools (DBpedia SPARQL Protocol and RDF Query Language (SPARQL)/lookup/schema exploration, Wikipedia-focused retrieval, and topical web search). We evaluate on 60 examples each from Event-QA and CMV (3 splits of 20), and report both mean end-to-end latency and per-example token costs. We evaluate GPT-4o and GPT-4o-mini under identical workflows and report accuracy and end-to-end latency. On EventQA, the best tool-augmented configuration improves accuracy (e.g., $47.5 \% \rightarrow 67.5 {\%}$ for GPT-4o) while increasing latency by orders of magnitude $(\sim 8 \mathrm{s} \rightarrow \sim 317 \mathrm{s}$ per example). On CMV, one-shot prompting is strongest (e.g., GPT-4o-mini achieves 75% accuracy at $\sim 6 \mathrm{s}$), and planning+search increases latency substantially without consistent gains. However, complex multitool orchestration exposes failure modes where the smaller model degrades. Overall, the findings highlight the need for taskspecific, cost-aware choices of both model size and agent/tooling complexity.
Paper
References (31)
Scroll for more · 19 remaining