arxiv:2504.00891

GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning

Published on Apr 1

· Submitted by

RyanLiu112 on Apr 4

Upvote

Authors:

Runze Liu ,

Kaiyan Zhang ,

Junqi Gao ,

Jiafei Lyu ,

Zhouyi Qian ,

Biqing Qi ,

Bowen Zhou

Abstract

Recent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However, current PRMs face three key challenges: (1) limited process supervision and generalization capabilities, (2) dependence on scalar value prediction without leveraging the generative abilities of LLMs, and (3) inability to scale the test-time compute of PRMs. In this work, we introduce GenPRM, a generative process reward model that performs explicit Chain-of-Thought (CoT) reasoning with code verification before providing judgment for each reasoning step. To obtain high-quality process supervision labels and rationale data, we propose Relative Progress Estimation (RPE) and a rationale synthesis framework that incorporates code verification. Experimental results on ProcessBench and several mathematical reasoning tasks show that GenPRM significantly outperforms prior PRMs with only 23K training data from MATH dataset. Through test-time scaling, a 1.5B GenPRM outperforms GPT-4o, and a 7B GenPRM surpasses Qwen2.5-Math-PRM-72B on ProcessBench. Additionally, GenPRM demonstrates strong abilities to serve as a critic model for policy model refinement. This work establishes a new paradigm for process supervision that bridges the gap between PRMs and critic models in LLMs. Our code, model, and data will be available in https://ryanliu112.github.io/GenPRM.

View arXiv page View PDF Project page GitHub repository Add to collection

Community

RyanLiu112

Paper author Paper submitter 1 day ago

We propose GenPRM, a strong generative process reward model with the following features:

reasoning with explicit CoT and code verfication before providing the process judgment;
improving Monte Carlo estimation and hard label with Relative Progress Estimation (RPE);
supporting GenPRM test-time scaling with majority voting;
supporting policy model test-time scaling with GenPRM as verifiers or critic models.

librarian-bot

about 8 hours ago

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

Your need to confirm your account before you can post a new comment.

· Sign up or log in to comment

Upvote

Models citing this paper 2

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2504.00891 in a Space README.md to link it from this page.