Researchers propose the Piggyback Hypothesis as an explanation for emergent misalignment, where fine-tuning on a narrow task causes unwanted behavior in unrelated prompts.
The paper argues that chat-template tokens can help transfer the fine-tuned behavior outside the original training domain, and tests this by perturbing those tokens.
The finding is relevant for AI safety work because it suggests small formatting details may influence how alignment failures generalize.