Raw data instead of ready-made tables
Models are increasingly acting as coding agents: they are tasked with writing scripts for data processing, building plots, computing metrics, and parsing engineering exports. But there is a class of tasks that has hardly been tested until now — working with physical-layer "raw material." This is precisely the niche filled by the EMRB (Electromagnetic Reasoning Benchmark), described in the arXiv preprint arXiv:2608.24086 (sections cs.AI, cs.CE, cs.SE).
The core idea is simple and therefore compelling. The input is only raw I/Q recordings, that is, quadrature samples of a signal as they arrive from a receiver. No preprocessed features, no labeled spectrograms, and certainly no tables with ready-made values. The quantity the question refers to must first be discovered by the model in the data — through code that it writes and runs itself.

How this differs from conventional benchmarks
In recent years, quite a few task sets have appeared in which models are asked to reason about radio signals. But most often the work has already been done for the model: features were extracted from the recording, the spectrum was computed, the noise level was estimated, and everything was compiled into a structured table. In such a setup, what remains is arithmetic and number matching — useful skills, but unrelated to radio physics.
The difference is fundamental. When the quantities are given in advance, a model's error at an early step goes unnoticed: an incorrect estimate of bandwidth or frequency simply never arises, because those values were provided in the problem statement. When the data is raw, any error at the signal reconnaissance stage drags down the entire subsequent calculation. That is why EMRB tests not knowledge of terminology, but the ability to build a chain: look at the samples, figure out what kind of signal is there, isolate the required characteristic, compute it correctly, and not lose track of the units.
How the benchmark is structured
The authors — Mingxu Zhang, Ying Sun, Yuhan Li, Yang Ji, Dazhong Shen, Ke Zhang, and Shan Huang — assembled 200 tasks. The tasks are divided into five difficulty levels and 27 question types: from trivial signal detection to designing an OFDM system. The source material is 11 signal types, each with verified ground truth prepared.
Five levels
The levels are arranged by increasing model autonomy. At the bottom are measurements of basic parameters: find the signal, estimate its characteristics. Above that — interpretation and comparison, then multi-step analysis, then diagnostics, and finally system design, where the model is required not to measure but to design a solution to given requirements.
27 question types and 11 signal types
Such a spread is needed so that the result does not depend on a single lucky or unlucky phrasing. Different signal types present different traps: in some cases noise gets in the way, in others it is spectral overlap, and in others one must handle sampling carefully. The number of question and signal types together forms a space in which it is hard to guess by chance.
Open data
All materials and code are published in a public GitHub repository, so the result can be independently verified. For a benchmark, this matters more than usual: if the tasks are generated rather than manually collected from field recordings, the question of ground-truth correctness becomes central — and openness is the only working answer here.
What the models showed
Fourteen language models were tested: proprietary ones, open-weight ones, and separately, models aimed at reasoning. The overall spread — from 24.1% to 78.9%. The range alone says a lot: the difference between the best and worst model here is measured not in percentage points but in multiples.
But something else is more interesting. The average result drops sharply as the level grows more complex: from 84.9% on basic measurements to 21.2% on system design. That is, models do reasonably well when they need to compute a measurable quantity, and almost fall apart when they need to assemble a working solution from those quantities.

This should be read as a diagnosis, not a ranking. The strength of current models is code execution and arithmetic on top of a known problem statement. The weakness is translating an engineering task into a language of measurable steps: figuring out which quantities are even needed, in what order to look for them, and how to check that what was found even resembles the truth. It is precisely this gap that makes the benchmark useful.
ReconPilot: reconnaissance, analysis, verification
The authors did not stop at measurement. They proposed ReconPilot — a structured approach in which the work is split into three stages: signal reconnaissance, targeted analysis, and self-verification. First, the model studies the recording and forms an idea of what it is dealing with. Then it solves the specific task. Then it returns to its result and checks it for consistency.
On three base models, the method adds between 3.8 and 17.6 points to the overall score. Improvement was achieved in 13 of the 15 tested "backbone + level" combinations. Note the wording: the gain is uneven, and in two cases there is none at all. This is a normal sign of an honest experiment — no universal cure was found, but the trend is consistent.
It is telling that what helps is precisely the explicit separation of phases. A model that is simply asked to "analyze the signal" tends to jump to calculations without making sure it has understood the data. Reconnaissance as a separate mandatory step forces it to look first and compute later.
What follows from this
For engineering practice, the conclusion is fairly direct: before trusting an agent with calculations based on real measurements, it is worth testing it on data without hints. A convenient interface and a smooth answer in chat say nothing about whether the model will survive an encounter with an unprepared export from a receiver.
There is also a more general idea that goes beyond radio. In any field where reasoning relies on raw measurements — hydroacoustics, vibration, telemetry — an agent must be able to first find the quantity in the data and then work with it. Tables with features hide this part of the work, and along with it, the errors are hidden too.
Limitations and what's next
The question cannot be called fully closed. 200 tasks is a decent volume, but not unlimited, and generation based on 11 signal types means that real field recordings with all their artifacts are still not represented there. Evaluating 14 models is a snapshot in time, not a verdict: the lineup of leaders in this field changes quickly.
Nevertheless, the formulation of the task looks right. If we want LLMs to work not with a retelling of data but with the data itself, we need to test them exactly where the hints end. EMRB does precisely this — and shows that the room for growth in signal reconnaissance is still very large.



