What Is DSPA
DSPA (Dynamic SAE Steering for Preference Alignment) is an inference-time preference alignment method. It does not update the base model’s weights, but uses preference data to steer generation.
The method requires preference triples. Based on them, DSPA builds a map linking prompt features to generation-steering features.
How DSPA Changes Generation
During decoding, the method modifies only the latents active for the current token. It leaves the other latents untouched.

Main steps:
- DSPA receives preference triples.
- It builds a conditional-difference map between prompt features and steering features.
- During decoding, it modifies the latents active for the token.
What the Experiments Showed
On Gemma-2-2B/9B and Qwen3-8B, the method improved results on MT-Bench. On AlpacaEval, DSPA was competitive, and it preserved accuracy on multiple-choice tasks.
With limited preference data, the method remained robust and could rival the two-stage RAHF-SCIT pipeline. DSPA required up to 4.47× fewer FLOPs for alignment.
The paper was accepted to the EMNLP 2026 Main Conference. The results apply to the models and tasks specified.
What Signals the Method Uses
An audit of the SAE features being modified showed that preference directions are primarily associated with discourse and stylistic signals.
The paper also presents a theory that refines the estimate of the conditional-difference map. It describes the conditions under which top-k ablation is justified.
When to Consider DSPA
The method may be a candidate for tasks that require preference alignment without updating the base model’s weights. It requires preference triples to work.
A practical criterion is to compare the task with results on MT-Bench, AlpacaEval, and multiple-choice tasks. When comparing with RAHF-SCIT, FLOPs during the alignment stage should be considered separately.



