Why trajectory is a bad unit of measurement
Modern inference-time reasoning methods all follow the same pattern: the model generates not a single answer but a whole pool of candidate chains, after which an external reward model picks the best one. The logic is simple — the more options there are, the higher the chance that one of them turns out to be a good one.
The weak spot lies elsewhere. Each candidate is treated as an indivisible whole: it is either accepted in full or thrown out in full. But a long reasoning chain is rarely homogeneous. The typical picture is a confident, correct start, a careful problem setup, a correctly chosen method, and then a breakdown: an arithmetic error, a loop, a substituted condition. Formally the entire chain is marked as a failure, even though its first half was perfectly workable.
The result is wasted computation. We pay for generating tokens that were correct, and then discard them along with the corrupted tail. It is precisely this inefficiency that the authors of Selective Regenerative Decoding point to, published as arXiv:2608.24338 (cs.AI) on August 25, 2026.

SRD: three branches instead of two
Selective Regenerative Decoding (SRD) proposes abandoning the binary choice. Instead of "accept or discard," each candidate is routed into one of three branches:
- discard — if the chain is useless from start to finish;
- keep — if it passes verification in full;
- refine — if there is a useful part, but the suffix has clearly degraded.
In the third case, the system does not rewrite the reasoning from scratch but preserves the high-quality prefix and regenerates only the corrupted section. This is closer to editing a draft than to writing a new text: you don't throw away a good introduction because of a failed ending — you fix what broke.
The authors separately emphasize that this kind of intervention does not require a larger donor model. No "smart teacher" pulling up a weak student — all the work happens within the computational resources already available. That is precisely what makes the method practical rather than a theoretical exercise.
Where the gain comes from
The most interesting thing about the work is not the heuristic but its justification. Under mild assumptions, SRD yields a provable 1.28–1.36× increase in sampling efficiency relative to classical rejection sampling, and the expected quality of the final trajectory turns out to be strictly higher, not merely "no worse."
The intuition here is clear without formulas. A candidate that has gone through suffix regeneration already contains paid-for computation — those prefix tokens that don't need to be generated again. The more such borderline candidates can be saved, the smaller the share of generated text that ends up in the trash. And since a larger candidate pool also means more "almost successful" chains, the gain does not remain constant — it scales along with the generation budget.

Where it was tested
The empirical part covers four different types of workload: MATH500 (math problems), GPQA Diamond (science-level questions), HotpotQA (multi-hop questions over documents), and AlpacaEval (instruction following and preference evaluation). Testing was conducted on several "generator + reward model" combinations so that the result would not depend on a single lucky pairing.
The reported results: SRD reaches the accuracy typically obtained with Best-of-N mode but consumes noticeably fewer generated tokens. And in compute-constrained scenarios, the method outperforms speculative rejection — that is, it wins precisely where every extra token is expensive.
What this changes in substance
The main shift the authors identify is methodological. Previously, selection happened at the level of the entire trajectory: the reward model assigned a score and decided the fate of the whole reasoning chain. SRD moves the intervention inside the trajectory, to the level of segments, and thereby opens up a region of trade-off between accuracy and computation that had until now been almost unexplored.
The practical takeaway for those building inference pipelines: the quality of reasoning and the cost of obtaining it are not necessarily rigidly linked quantities. Some of the "extra" computation can be recovered if you stop treating the model's draft as a monolith.
What remains off-screen
There are plenty of open questions. The key one is exactly how the boundary between a healthy prefix and a degraded suffix is determined: how many candidates even make it into the third branch depends on this criterion. Then there is the cost of regeneration itself (it is not free), sensitivity to the choice of reward model, and the transferability of the conclusions beyond the four benchmarks used.
The paper is 20 pages long and was submitted to ARR, meaning it is still a first-version preprint rather than a peer-reviewed result. But the direction looks promising: instead of generating more and selecting more harshly, one can make more careful use of what the model has already come up with.




