A new arXiv paper presents turn-averaged sparse autoencoders for feature discovery and long-context attribution. The approach is aimed at understanding model behavior across multi-turn interactions rather than isolated prompts.
That matters because many real LLM deployments involve conversations, agents and documents that stretch across long contexts. Interpretability methods need to explain patterns over time, not only single-token decisions.
The paper adds to the growing toolkit for mechanistic interpretability as researchers try to make large models easier to inspect and debug.