How to Reuse Valuable Generations in GRPO Training

25 September 202615 views

The article presents a mechanism that selects individual generations for retraining based on advantage magnitude and discards overly stale data, supplementing it with fresh examples. Across three Qwen3-Base scales, the method outperformed GRPO and standard replay on math benchmarks, while also improving the balance between accuracy and token usage.

How to Reuse Valuable Generations in GRPO Training

What exactly is reused

In standard GRPO, each rollout is used in one gradient update and then discarded. The method described in the paper changes this cycle: it retains individual rollouts and returns selected ones to training.

The buffer stores individual generations, not entire groups. Replay prioritizes them by advantage magnitude: rollouts with higher values are reused.

A practical criterion: this is about reusing individual rollouts, not groups.

How replay works

Each batch is formed in two stages:

  1. Fresh on-policy rollouts are included.
  2. Separately selected rollouts from the buffer are added.

Replay items are selected based on the advantage magnitude of each rollout. The reported facts do not specify the share of replay items in a batch.

The key difference is that fresh rollouts remain in the batch, while the buffer supplements them with selected generations.

How staleness is limited

The authors note a problem with naive replay: LLM policies change quickly with each gradient step. Stored rollouts can become stale and destabilize training.

The method removes any rollout from the buffer that is older than tau_max training steps. This sets a retention limit, but does not eliminate reuse itself.

The control criterion here is the age of an individual rollout in training steps, not its membership in a group.

What the evaluation showed

The method was tested at three Qwen3-Base scales and on five math benchmarks. It outperformed GRPO and naive replay baselines.

  • At every scale, the gain in the average across the five benchmarks was positive.
  • At the 4B scale, the gain reached +1.66 pp.
  • On AES, a joint measure of accuracy and token efficiency, the method was the only option to outperform GRPO at every scale.

The practical selection criterion depends on the target evaluation: the results cover accuracy, token efficiency, and comparison with two baselines.

Frequently asked questions

How to Reuse Valuable Generations in GRPO Training