Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning

The Process Reward Model (PRM) plays a crucial role in mathematical reasoning tasks, requiring high-quality supervised process data. However, we observe that reasoning steps generated by Large Language Models (LLMs) often fail to exhibit strictly incremental information, leading to redundancy that can hinder effective reasoning. To address this issue, we propose CFPRM, a simple yet effective co…

Paper

Similar papers

© 2026 NYSGPT2525 LLC