DSPA adjusts language model responses based on preferences during generation

26 September 202613 views

The method maps query features to SAE directions and applies them to active token representations without retraining the base model. Tests on Gemma-2 and Qwen3 showed improved results on MT-Bench; with a small amount of preference data, DSPA requires up to 4.47 times less compute for alignment than the RAHF-SCIT pipeline.

DSPA adjusts language model responses based on preferences during generation

What Is DSPA

DSPA (Dynamic SAE Steering for Preference Alignment) is an inference-time preference alignment method. It does not update the base model’s weights, but uses preference data to steer generation.

The method requires preference triples. Based on them, DSPA builds a map linking prompt features to generation-steering features.

How DSPA Changes Generation

During decoding, the method modifies only the latents active for the current token. It leaves the other latents untouched.

Main steps:

  1. DSPA receives preference triples.
  2. It builds a conditional-difference map between prompt features and steering features.
  3. During decoding, it modifies the latents active for the token.

What the Experiments Showed

On Gemma-2-2B/9B and Qwen3-8B, the method improved results on MT-Bench. On AlpacaEval, DSPA was competitive, and it preserved accuracy on multiple-choice tasks.

With limited preference data, the method remained robust and could rival the two-stage RAHF-SCIT pipeline. DSPA required up to 4.47× fewer FLOPs for alignment.

The paper was accepted to the EMNLP 2026 Main Conference. The results apply to the models and tasks specified.

What Signals the Method Uses

An audit of the SAE features being modified showed that preference directions are primarily associated with discourse and stylistic signals.

The paper also presents a theory that refines the estimate of the conditional-difference map. It describes the conditions under which top-k ablation is justified.

When to Consider DSPA

The method may be a candidate for tasks that require preference alignment without updating the base model’s weights. It requires preference triples to work.

A practical criterion is to compare the task with results on MT-Bench, AlpacaEval, and multiple-choice tasks. When comparing with RAHF-SCIT, FLOPs during the alignment stage should be considered separately.

Frequently asked questions

Related materials

All materials
DSPA adjusts language model responses based on preferences during generation