What the simulation evaluates
TRACE is a diagnostic environment for digital advertising. It covers 12 root causes and detailed segment attribution. In each episode, the agent investigates data using Python and SQL. It then identifies the root cause and, where applicable, the affected segment.
TRACE is described in a paper by Rui Sun, Zhan Shi, and Bing He, submitted to arXiv on September 9, 2026. The paper ID is arXiv:2609.10315.

How the evaluation is structured
The evaluation is built around an intervention and the observations it triggers:
- An intervention is selected and introduced into a controlled simulator.
- The simulator generates the corresponding observations.
- The agent investigates the data and looks for the cause of the failure.
- The hidden intervention serves as the ground-truth label and makes it possible to calculate an objective reward.
The agent is not given a ready-made explanation. It must piece together noisy, mixed, and distributed evidence. This means the simulation evaluates causal inference, not just recognition of a predefined answer.
How to read the results
The authors evaluated the models on a held-out test set of 235 episodes. The strongest prompting baseline was Claude Opus 5, with a FullAttr@1 score of 0.686.
| Model or stage | Result |
|---|---|
| Qwen3.5-35B-A3B after supervised fine-tuning | 0.159 → 0.637 |
| The same model after RL with synthesized rewards | 0.757 |
| Qwen3.5-122B-A10B in a prompting baseline | Below the 0.757 result |
The score of 0.757 exceeded all evaluated prompting baselines, including closed-source frontier models. The authors also report that the resulting policy used substantially fewer tool calls than the prompting version of the 35B base model.
What the results show about agent training
The authors consider access to a scalable objective training signal a potentially more important constraint than model size alone. This is a conclusion based on the described task, not a universal rule for all AI systems.
A simulation-based evaluation can make ambiguous diagnostic reasoning tasks suitable for scalable reinforcement learning. The key condition is being able to compare the agent’s conclusion with a hidden intervention and calculate an objective reward.
When this approach is appropriate
Simulation is suitable for evaluating causal inference if the task allows you to:
- define an intervention and introduce it into a controlled environment;
- generate observations associated with it;
- retain the intervention as a hidden ground-truth label;
- evaluate both the root cause and the affected segment, where applicable.
If the task does not allow this kind of comparison with a ground-truth label, the described objective reward mechanism is not available. A practical selection criterion is whether the agent’s conclusion can be checked against a hidden cause without revealing it during analysis.



