RefuteBench 2.0 -- Agentic Benchmark for Dynamic Evaluation of LLM Responses to Refutation Instruction

While large language models (LLMs) can leverage user feedback in multi-turn interactions, evaluating their ability to incorporate refutation feedback remains challenging. We introduce RefuteBench 2.0, an evaluation framework using LLM agents as refuters and evaluators to assess how models handle contradictory input. The framework includes transient and persistent refutation instructions with varying validity periods to reflect real-world interaction patterns. Meta-evaluation shows our approach generates human-like refutations and achieves high correlation with human evaluators. Experiments across various LLMs reveal a critical limitation: while models can satisfy immediate refutation requests, they fail to retain refutation information over extended dialogues. Notably, increasing refutations progressively degrades performance on the original task. Attention analysis confirms that current LLMs struggle with information retention and correct utilization in long-context scenarios, highlighting a fundamental weakness in their multi-turn reasoning capabilities.

Paper

References (43)

Scroll for more · 31 remaining

Similar papers

© 2026 NYSGPT2525 LLC