Simulation for Testing Causal Inferences About Advertising Failures

29 September 202612 views

In TRACE, hidden interventions in the digital advertising simulator make it possible to automatically assess whether an agent correctly identified the source of a failure and the affected segment, even though it must analyze noisy data using Python and SQL. After reinforcement learning on synthesized signals, the Qwen3.5-35B-A3B model scored 0.757 on FullAttr@1 across 235 test episodes—higher than all tested prompted variants, while using fewer tool calls than the baseline model.

Simulation for Testing Causal Inferences About Advertising Failures

What the simulation evaluates

TRACE is a diagnostic environment for digital advertising. It covers 12 root causes and detailed segment attribution. In each episode, the agent investigates data using Python and SQL. It then identifies the root cause and, where applicable, the affected segment.

TRACE is described in a paper by Rui Sun, Zhan Shi, and Bing He, submitted to arXiv on September 9, 2026. The paper ID is arXiv:2609.10315.

How the evaluation is structured

The evaluation is built around an intervention and the observations it triggers:

  • An intervention is selected and introduced into a controlled simulator.
  • The simulator generates the corresponding observations.
  • The agent investigates the data and looks for the cause of the failure.
  • The hidden intervention serves as the ground-truth label and makes it possible to calculate an objective reward.

The agent is not given a ready-made explanation. It must piece together noisy, mixed, and distributed evidence. This means the simulation evaluates causal inference, not just recognition of a predefined answer.

How to read the results

The authors evaluated the models on a held-out test set of 235 episodes. The strongest prompting baseline was Claude Opus 5, with a FullAttr@1 score of 0.686.

Model or stageResult
Qwen3.5-35B-A3B after supervised fine-tuning0.159 → 0.637
The same model after RL with synthesized rewards0.757
Qwen3.5-122B-A10B in a prompting baselineBelow the 0.757 result

The score of 0.757 exceeded all evaluated prompting baselines, including closed-source frontier models. The authors also report that the resulting policy used substantially fewer tool calls than the prompting version of the 35B base model.

What the results show about agent training

The authors consider access to a scalable objective training signal a potentially more important constraint than model size alone. This is a conclusion based on the described task, not a universal rule for all AI systems.

A simulation-based evaluation can make ambiguous diagnostic reasoning tasks suitable for scalable reinforcement learning. The key condition is being able to compare the agent’s conclusion with a hidden intervention and calculate an objective reward.

When this approach is appropriate

Simulation is suitable for evaluating causal inference if the task allows you to:

  • define an intervention and introduce it into a controlled environment;
  • generate observations associated with it;
  • retain the intervention as a hidden ground-truth label;
  • evaluate both the root cause and the affected segment, where applicable.

If the task does not allow this kind of comparison with a ground-truth label, the described objective reward mechanism is not available. A practical selection criterion is whether the agent’s conclusion can be checked against a hidden cause without revealing it during analysis.

Frequently asked questions

Simulation for Testing Causal Inferences About Advertising Failures