Full fine-tuning is no longer a reasonable option
Foundation models for electroencephalography are trained without labels — on contrastive tasks or on reconstructing masked fragments of the signal. The logic is this: the model picks up general patterns of rhythms, artifacts, and transitions between states, and then it is fine-tuned for a narrow task like detecting epileptic discharges. In theory, this yields representations that transfer across clinics and equipment. In practice, it is more complicated: clinical recording archives are too dissimilar from one another, and transfer quality drops noticeably.
The second problem is economics. Classic fine-tuning updates all weights, which means it requires serious computational resources for every new task. In a department without a GPU cluster, such a scenario simply cannot run. Hence the question the paper arXiv:2608.24727 is built around: can you get by with a small fraction of parameters without losing quality?

Nine percent: what exactly gets tuned
The authors of the preprint are Meghal Dani and Stefanie Liebe. They test a regime in which roughly 9% of the weights remain trainable while the rest of the model is frozen. Adaptation proceeds in a self-supervised manner: the model is not simply fitted to class labels but aligns its internal representations with the target task. The text was posted on August 25, 2026 (version v1), assigned to the cs.LG and cs.AI sections, DOI — 10.48550/arXiv.2608.24727. The code is open for reproduction.
The point is not that the result is a "trimmed-down" model. It is a different mode of operation: the heavy pretrained encoder stays unchanged, and a compact add-on is tuned for each new task. For a clinic, this separation is fundamental — expensive pretraining is done once, cheap adaptation is repeated as many times as needed.
Two models with different pretraining objectives
To keep the result from depending on a single lucky architecture, the experiment involves two SOTA models: BIOT with a contrastive objective and CBraMod, trained to reconstruct masked regions. If the technique works in both cases, it is not a coincidence but a property of the approach.
Three clinical datasets and two evaluation regimes
Evaluation is performed on TUAB (anomaly detection), TUEV (event classification), and CHB-MIT (seizure detection). Measurements are taken twice — in-distribution and out-of-distribution, meaning they test not only on "their own" data but also on a shifted distribution.
What the results showed
- Compared with linear probing, self-supervised adaptation yields a consistent gain — up to 20-fold in AUCPR. This is not "slightly better" but a difference of an order of magnitude, and it is especially noticeable where a rare class matters more than overall accuracy.
- With a fixed computational budget, the peak is reached with just 20–50% of the available unlabeled recordings. The rest of the corpus adds no further gain.
- If the total number of signal windows is fixed, the result does not depend on how many patients are included. So what works is not the "uniqueness of people" but the temporal diversity of fragments: different states, modes, and artifacts within a recording.
The last point upends the usual logic of data collection. Intuitively, it seems that the more different patients, the more reliable. Here, it turns out that under a strict limit on the number of windows, you can gather data from fewer people — it is only important to monitor the diversity of their signal over time.

Why this changes the calculus for clinics
Until now, the adoption of EEG foundation models has been hampered not by the quality of representations but by budget: to get value from the model, you had to be able to afford full fine-tuning and a large labeled corpus. The work shows that this barrier can be lowered from two sides at once — in computation and in data. Nine percent of the weights and half of the unlabeled archive instead of one hundred percent and the entire volume.
For a hospital, this means a more realistic scenario: the model is pretrained centrally, and on site it is tuned for a specific task on modest hardware. Less dependence on who collected more recordings, and lower costs for storing and processing unnecessary hours of signal.
What to keep in mind
The result should not be read as "nine percent is always enough." This is a specific regime from a specific experiment, not a universal recipe for any architecture and any dataset. The authors themselves note: generalization of EEG models remains limited, especially on heterogeneous clinical corpora. Parameter-efficient adaptation does not eliminate the transfer problem — it merely lowers the cost of entry into it.
The practical takeaway is simpler than it seems: the conversation shifts from "should we even bother with foundation models for EEG" to "how do we deploy them in the department." And that is already a matter of engineering and process organization, not just machine learning.




