Coding agents spend a meaningful share of their budgets repeating work, according to an analysis of 1,200 Claude Code and Mini-SWE-Agent trajectories on SWE-bench Verified. The researchers identified three recurring patterns: retrieving information already covered by earlier results, generating similar scripts, and rerunning tests unnecessarily.
Those behaviors appeared in 79% to 98% of tasks and represented as much as 22.75% of task cost. The team then tested mitigation strategies across more than 10,000 held-out trajectories from SWE-bench Verified and Pro. Some seemingly obvious fixes backfired: structure-aware retrieval added overhead and altered delegation, increasing total cost by as much as 28.14% in some configurations.
Skills synthesized by agents tended to encode low-level advice tied to a particular trace, limiting reuse. By contrast, developer-designed skills gave broader guidance that transferred across tasks and cut costs by up to 41.73%, roughly twice the best reduction from agent-written skills. These are benchmark costs rather than savings guaranteed for every repository or provider. The findings suggest teams should inspect complete action traces, not only token prices, and turn recurring human lessons into concise, task-independent operating rules.