What Remains Human: How to Build Ethical LLM Work in Science

15 September 202624 views

A new paper proposes a formal language for describing which parts of scientific reasoning can be delegated to a model and which cannot. The key criterion for ethicality is not the extent of machine involvement, but the quality of verification and the presence of a specific person accountable for the result.

What Remains Human: How to Build Ethical LLM Work in Science

The tool is already in the lab — the question is different

Language models are no longer an exotic novelty in research work. Literature synthesis, hypothesis drafts, data processing code, checking formal derivations — all of this has already become part of everyday practice. But along with convenience comes a question that cannot be resolved through settings and regulations: if part of the reasoning was done by a machine, under what conditions can the result still be called knowledge rather than merely plausible text?

This is precisely how Kalin Stoyanov frames the problem in the paper "Ethical LLM-Assisted Research" (arXiv:2608.23644, August 24, 2026). This is not another call to "use AI carefully," but an attempt to build a normative and conceptual framework — that is, a language in which the delegation of cognitive work can be discussed rigorously.

The key premise: scientific reasoning is a distributed process. Contributions to it can come from various sources, human and machine, in any proportion. But responsibility for what ultimately enters the scientific record remains human. This means the question "how much AI is acceptable in a paper" is formulated incorrectly. The right question is different: what exactly is a person obligated to do so that the result remains legitimate.

Five variables instead of one vague "AI involvement"

Usually the debate boils down to a percentage: how much was written by a human, how much by the model. The framework proposes breaking this down into five separate constructs. Each of them can vary independently of the others.

  • Origin of content (O(g)) — who or what generated a specific claim.
  • Completeness of human verification (V(g)) — whether the verification was carried through to the end by a living researcher.
  • Assignment of responsibility (R(g)) — to whom the burden of the decision to accept a claim is attributed.
  • Accountable human ownership (M(g)) — who owns the claim in the sense of being ready to defend it and answer for it.
  • Epistemic outcome (E(g)) — what the verification yields in terms of knowledge gain.

The distinction looks almost pedantic, but it closes a loophole that everyone uses. Origin, verification, outcome, and responsibility are different things, and one cannot be derived from another. Text could have been generated by a machine, and a human will still be responsible for it — and that is fine. The reverse also happens: formally everything was written by the author, but the verification was fictitious, and knowledge ultimately does not grow.

Why you can't get away with "I checked everything"

The phrase "I read it and I agree" is not verification. It does not record what exactly was checked, on what basis, and where the boundary lies between the confirmed and the taken on faith. The construct V(g) specifically requires that verification be carried through, not merely declared. This matters because LLM-assisted reasoning is outwardly almost indistinguishable from ordinary reasoning: the text is smooth, the arguments are coherent, the references look real.

The boundary runs along verification, not along the volume of machine contribution

The central thesis of the paper is this: the ethical boundary is determined by adequate verification and accountable human ownership, not by the degree of machine involvement. That is, the very fact of delegation is ethically neutral. It becomes a problem only when verification is formal and responsibility is blurred.

This overturns the usual logic. Arguing about "what percentage" was written by the model is pointless: a single line generated and blindly inserted into a conclusion can ruin the result, and extensive machine assistance with honest verification breaks nothing. What is dangerous is not the volume of delegation but its opacity.

Epistemic audit: a log that makes reasoning verifiable

From this logic grows a practical tool — the epistemic audit. It is a structured record of what was delegated, how it was verified, where the content came from, and who bears responsibility. In essence — a log of the provenance of ideas, not a list of tools used.

The value of such a record is that it makes reasoning available for external scrutiny. A reader, reviewer, or co-author can see where the machine's guess ends and a confirmed fact begins. This is not a report for regulators but a working artifact: it also helps the author notice that somewhere they accepted a claim without verification simply because it sounded convincing.

What this looks like in daily work

The framework can be reduced to a few actions that require neither a new position in the lab nor special software.

  • Keep a delegation log: what exactly you handed over to the model — literature search, a draft argument, code, formal verification.
  • Mark provenance: which fragment came from the machine and which from you.
  • Separate verification from agreement: record what exactly you verified and by what method.
  • Name the responsible person by name, not "the group of authors." It doesn't matter whether it was ChatGPT or another assistant — a human is responsible for the claim.
  • Check the most convincing things first: a smoothly written paragraph is not a sign of correctness, and sometimes the opposite.

Where it most often breaks

Three typical failure points. The first — references and citations: they look plausible but require manual confirmation. The second — numerical and statistical claims, which are easy to take for granted. The third — "self-evident" generalizations that no one checks because they seem banal.

Limits of applicability and contested points

The framework is formal, which means it works with the consequences of adopted definitions rather than with living practice. The question of how to measure "completeness of verification" remains open — it is a construct, not a number. It is also not entirely clear how to embed the audit into peer review: journals do not yet require delegation logs, and it is not certain they will.

There is also the risk of the idea turning into a ritual. Any reporting tends to become bureaucracy: if the audit is done for show, it loses meaning in exactly the same way a fictitious verification would. The usefulness of the tool rests on the honesty of the person keeping it.

What ultimately remains for the human

Much can be handed over to the machine: draft work, exploring options, the routine of searching, the initial assembly of text. What is not delegated is something else — the decision about what counts as established, and the readiness to answer for it. Stoyanov's framework provides a formal vocabulary for this: it allows distinguishing responsible cognitive delegation from a simple transfer of responsibility or from its quiet ignoring.

In effect, this is a return to the old norm of scientific work, only with a new tool in hand. Knowledge has always been the result of verification, not of assertion. The emergence of LLMs has changed nothing in this — it has only made the difference between verification and its imitation more noticeable.

Frequently asked questions

Related materials

All materials
What Remains Human: How to Build Ethical LLM Work in Science