Why Block Summaries May Not Be Enough
Block-level residual routing makes learnable aggregation of residual connections practical: the router relies on block summaries. Each summary compresses the ordered sequence of attention and MLP updates into a single cumulative vector.
- The summary preserves the block’s total contribution.
- It does not describe the trajectory of updates within the block.
Practical criterion: if a rough signal of how the residual stream moves within a block matters, a single cumulative summary may not be enough.
How HAARES Adds an Intra-Block Signal
HAARES is a residual-basis router proposed in the paper by Kehan Wang, arXiv:2606.06564. It preserves the cumulative block source and adds a detail basis with a half-split.
- The detail basis is computed as the difference between residual stream updates in the first and second halves of the block.
- The basis is RMS-aligned and updated online.
- The router gets coarse information about the intra-block trajectory without dense routing at the sublayer level.

Practical criterion: HAARES combines a cumulative source with one additional detail signal, rather than routing each sublayer separately.
Where the Effect Appeared in Experiments
The experiments covered OpenWebText, cross-domain character-level benchmarks, and OpenWebText with BPE tokenization. The result depended on model depth.
| Condition | Observation |
|---|---|
| Shallow models | Gains are small or ambiguous |
| 48-layer models | The effect is most consistent |
| 201M configuration with 48 layers | HAARES outperforms Block AttnRes across all three seeds |
| 453M probe with two seeds | The result points in the same direction |
Ablations show that the effect is not explained by source duplication, random signed details, or fixed offsets for detail sources. It also cannot be explained solely by changing the number of blocks.
Practical criterion: the results provide a stronger case for the method in 48-layer models than in shallow models.
What Costs to Consider
The cost analysis notes low additional FLOPs, but not zero time overhead. The method adds memory and routing costs.
- The relative arithmetic cost is amortized as model width increases.
- Earlier convergence can reduce the time to reach a target result.
Practical criterion: HAARES is worth evaluating not only in terms of FLOPs, but also memory, routing, and convergence time.



