Researchers report that steering vectors can transfer across some independently trained language models, suggesting that shared internal geometry can have functional consequences.

The study evaluates five open-weight models across parameter scales from 0.8 billion to 8 billion and two architectural lineages. The authors trained sparse autoencoders across 15 semantic domains and tested alignment across directed model pairs. They found a suggestive break near 1.7 billion parameters: larger models showed much stronger cross-model feature validation, while alignment degraded sharply at 0.8 billion.

The practical result is that concept directions from one model can sometimes steer another model’s behavior. The paper reports a 71 percent win rate for cross-model steering vectors across supervised concepts, compared with 68 percent for same-model native vectors in the reported setting.

The finding is conditional, not universal. It depends on model scale, architecture, and the concepts tested. Still, it matters for interpretability because it suggests some behavioral controls may not be locked to one model’s internal representation.