A protocol no one follows to the letter
A lab protocol is rarely followed literally. Between publication and the actual experiment there is almost always adaptation: a different cell line, a substituted reagent, fewer samples, someone else's instrument, a trimmed budget. For a person at the wet bench this is routine, but a treacherous kind of routine — any edit drags a chain of consequences behind it. Change the reagent volume at an early step, and you have to remember how that will play out at the wash stage and how it squares with what was already done before. It is precisely this ability — to hold the entire protocol in your head and modify it meaningfully — that the new task set tests.
The study came out under the title BenchBench-Protocol in the arXiv preprint arXiv:2608.23898v1 (cs.AI), submitted on August 24, 2026. The authors are Aditya Sivakumar, Ashu Singhal, Nicholas Larus-Stone, and Nithin Parsan. It runs 21 pages with 15 figures, so the material is more methodological than promotional.

How 149 tasks were assembled from draft edits
The mechanics of BenchBench-Protocol are unusual. The authors did not sit down to write questions — they took an existing trail of work. Every published protocol had a version that a scientist had edited for their specific experiment. The difference between those two documents became the task: the model is shown the original protocol and the conditions of the new experiment, and it has to propose a correct modification.
An important detail follows directly from this approach: the task does not have an abstract "correct answer" but a weighted rubric. That is because the edit made by a real researcher is known — it can be broken down into elements and scored by how closely the model's proposal matches the real decision in meaning rather than in wording. This removes the perennial problem of biomedical tests, where a model can give a reasonable answer that does not match the reference and score zero.
The foundation turned out to be solid: 96 original protocols from nine areas of wet-lab biology. Only those tasks that passed expert review with high ratings made it into the final set — the rest were filtered out. The result is those very 149 assignments.
Why this is not just another set of expert-written questions
The authors spell out the context separately. In recent years biomedical benchmarks have indeed shifted toward open-ended tasks scored by rubric — that is progress. But the tasks themselves are still most often invented by experts: they sit down and write questions based on their own experience. The limitation of this approach is that the expert answers a question they formulated themselves, implicitly tuning it to their own sense of what is difficult.
Here the logic is reversed. A task arises where a real person has already run into a real problem and spent time solving it. This does not guarantee perfectly clean data, but it does guarantee that the problem was not invented for the sake of a nice phrasing.
Who managed and who did not
Nine models took part in the testing — both closed and open. The spread turned out to be noticeable, though not dramatic.
The best result came from Claude Opus 5 — 59.2% of the normalized rubric score. The other models fell within a range of 34.1% to 47.1%. Put simply, no one reached the leader, but there is no one-and-a-half-times collapse either: all the models can do something, they just do it differently and not always all the way through.
A separate line in the report is the saturation check. The authors gave the models ten attempts each and picked the best one (best-of-10). Even under such a generous scheme the benchmark did not "collapse": the ceiling was not reached. For an evaluation this is good news — it means the task set will not go stale after the next generation of models arrives.

What follows from this in practice
The first conclusion is practical and slightly sobering. A language model cannot yet be fully trusted with protocol adaptation. Even the best result of 59.2% means that roughly four times out of ten the proposal requires human intervention. The model works as a draft, as a generator of options, as a way to quickly check what you may have overlooked. But it does not bear final responsibility for the test tube.
The second conclusion is methodological, and it is more interesting than the first. The authors see their work not only as an evaluation of reasoning in the wet lab but also as proof of a more general idea: real experiments are an underrated source of benchmark tasks. If a lab already has a history of edits, then it also has material for testing models — material, moreover, that cannot be invented at a desk.
The third conclusion concerns where to go next. The gap between 59.2% and 47.1% does not look like a chasm given the difficulty of the task itself. It is more of a hint that the bottleneck is not the volume of a model's knowledge about biology but its ability to build a chain: to account for previous decisions and predict where they will lead in the following steps. It is this quality, rather than fact memorization, that the set of 149 edits is devoted to.



