Thanks for your prompt reply. We are glad that we have addressed your other major concerns about **performance on self-consistency** and **missing comparisons**. We now would like to clarify the other two arguments.
> Inference Time
Thanks for bringing the inference time up. Actually, “combining DBS with SC adds no time” is the response to your concern about “Methods should directly compare rather than add”. DBS can cache all finished reasoning paths during inference. Thus, users can directly apply self-consistency on the DBS-generated reasoning paths, which makes DBS+SC add no time compared to DBS alone. \
As for your concern on DBS increasing inference time, we admit that applying DBS indeed increases inference time and may be challenging to be deployed as a delay-sensitive system in the near future. Yet, being able to be used directly in the real-world systems is *one of the many forms of metric* to evaluate the scientific contributions. We would like to highlight again that although accurate reasoning is more time-consuming, it catalyzes stronger models [1,2,3,4]. The cited works all use *“time-consuming”* sampling methods, which synthesize accurate reasoning paths, and finally breeds more powerful models. Moreover, as we have discussed in the response to Reviewer 2vf7, we compare the DBS inference time with self-consistency, where the additional time cost is primarily introduced by the small verifier, and verifier inference time is 0.1x of LLM inference time. Ultimately, in terms of future impacts, we believe that accuracy is a critical factor for stronger models. \
Thanks again for pointing out the limitation from the immediate use perspective, we will add more discussions in terms of inference speed in the future revision.
> Big Bench Hard
Thanks for the suggestion. We respectfully suggest that BBH may not be indispensable for work focusing on reasoning ability. As you pointed out, “reasoning capability is one of their focuses.” Given the multitude of abilities involved, including natural language understanding, it can be challenging to evaluate **reasoning ability**. Moreover, in recent works focusing on improving reasoning ability [5,6,7,8,9,10], BBH is not necessarily evaluated to demonstrate the effectiveness of reasoning. We are happy to discuss more if you could possible consider elaborating more.
Thanks again for your valuable suggestions and for being willing to discuss, let us know if you have any other questions.
[1] Kexun Zhang, et al. Algo: Synthesizing algorithmic programs with generated oracle verifiers, NeurIPS 2024.
[2] Trieu H. Trinh, et al. Solving olympiad geometry without human demonstrations, Nature, 2024.
[3] Jiaxin Huang, et al. Large Language Models Can Self-Improve. EMNLP 2023
[4] Zelikman Eric, et al. Star: Bootstrapping reasoning with reasoning. NeurIPS 2022
[5] Ning Miao, et al. SelfCheck: Using LLMs to Zero-Shot Check Their Own Step-by-Step Reasoning. ICLR 2024
[6] Terufumi Morishita, et al. Learning Deductive Reasoning from Synthetic Corpus Based on Formal Logic, PMLR 2023
[7] Yuxi Xie, et al. Self-evaluation guided beam search for reasoning, NeurIPS 2024.
[8] Zhan Ling, et al. Deductive verification of chain-of-thought reasoning, NeurIPS 2024.
[9] Shibo Hao, et al. Reasoning with language model is planning with world model, EMNLP 2023.
[10] Xidong Feng, et al. Alphazero-like tree-search can guide large language model decoding and training, 2023