Active Weights Instead of a Full Matrix: INT8 Decoding of a Spiking Model on CPU

27 September 202624 views

The article describes a C++ implementation for a language model with binary spike gating: sparse projections are processed in INT8, and computations account only for active weights. In a single-threaded test, the early INT8 version achieved 23.31 tokens/s versus 9.82 for FP32 and reduced memory usage for weights from 3355.2 to 1087.4 MiB; switching dense projections to INT4 reduced decoding throughput.

Active Weights Instead of a Full Matrix: INT8 Decoding of a Spiking Model on CPU

How spikes change weight processing

In the article “Spike-Aware INT8 Execution for Spiking Language Models on Commodity CPUs” Ting Liu describes decoding a spiking language model on a CPU.

  • Binary spike activations make it possible to read only the active weight columns.
  • Multiplications are replaced with sums of weights.

The practical point of this approach is to avoid processing columns that did not activate. The gains depend on activation sparsity, not just the weight format.

How the INT8 path works

The C++ implementation was tested on a spike-gated language model with 874M parameters. It uses different modes for sparse and dense projections.

  • Sparse projections store INT8 weights in column-major format.
  • Accumulation is performed using integers, with the scale applied once per output channel.
  • Dense projections retain row-major access and FP32 activations.

This is a mixed scheme, not a conversion of all model operations to INT8. The selection criterion is whether to retain FP32 in dense projections in favor of a sparse INT8 path.

What speed and storage metrics are reported

An early checkpoint was compared using a single thread. In this test, INT8 showed higher decoding throughput and lower weight storage.

Early checkpoint, single threadDecoding speedWeight storage
FP329.82 tokens/s3355.2 MiB
INT823.31 tokens/s1087.4 MiB

Separate results are reported for the final checkpoint on an AMD Ryzen 7 5800X. The conditions differ from the early-checkpoint comparison.

Final checkpoint INT8 modeMetric
Decoding, single thread22.63 tokens/s
Decoding, four threads47.90 tokens/s
Prefill of a 512-token sequence, eight threads94.68 tokens/s

These metrics cannot be directly combined into a single comparison table: they relate to different checkpoints and modes. Both the checkpoint version and the number of threads matter when evaluating throughput.

What the INT4 variant changes

The INT4 variant for dense projections reduces storage by a further 17.4%. Decoding throughput decreases by 46.6%.

The revised version of the article retracts invalid results for pure INT4. Therefore, the reported trade-off applies to INT4 in dense projections, not to a fully INT4 model.

The selection criterion is prioritizing storage over decoding speed. Pure INT4 results cannot be used as evidence of performance.

What is known about energy consumption

A separate study of the output head on ARM recorded higher energy figures for two candidate-checking configurations over a truncated decoding window.

The revised version of the article also adds numerical validation and a wall-power energy-consumption study on ARM. CPU speed test data alone is not enough to draw conclusions about energy savings.

Frequently asked questions

Related materials

All materials
Active Weights Instead of a Full Matrix: INT8 Decoding of a Spiking Model on CPU