Over-Reasoning and Redundant Calculation of Large Language Models

Large language models (LLMs) can solve problems step-by-step.While this chain-of-thought (CoT) reasoning boosts LLMs’ performance, it is unclear if LLMs know when to use CoT and whether those CoT are always necessary to answer the question. This paper shows that LLMs tend to generate redundant calculations and reasoning on a manually constructed math QA dataset, GSM8K-Zero.GSM8K-Zero is constructed such that the questions can be answered without any calculations, but LLMs, including Llama-2 models and Claude-2, tend to generate lengthy and unnecessary calculations to answer the questions.We also conduct experiments to explain why LLMs generate redundant calculations and reasonings.

Paper

References (19)

10. Model card and evaluations for claude models2023 · Anthropic
112023. Stanford alpaca: An instruction-following llama modelgithub
12This is highly unlikely to happenB More Information about GSM8K-Zero

Scroll for more · 7 remaining

Similar papers

© 2026 NYSGPT2525 LLC