Improvements in reasoning and coding can come at the expense of clear prose, according to Anthropic fine-tuning engineer Jackson Kernion. He says newer Claude models learned a dense style partly because reinforcement learning rewarded explanations that other language models could follow, not only text that reads naturally to people.
A model has far more working memory than a human reader and can track details through tightly packed explanations. Training on math, code and technical material therefore encourages output optimized for that audience. Kernion describes the result as overly dense information dumps and says teams must actively add rewards for simple, human-readable explanations to counter the tendency.
The account helps explain why a higher-scoring model may feel worse for editing, storytelling or ordinary communication. Capability is not one scale: a reward structure can improve performance on tasks with easily checked answers while weakening style, pacing and selection of detail. Users who choose a model by benchmark rank alone may therefore get a poor fit for writing work.
Kernion says Anthropic achieved a better balance with Opus 5.5 and that he has not been as satisfied with a model’s writing since Opus 4.6. He stops short of saying the new release surpasses the older one, and describes natural writing quality as an ongoing training problem.