The maintainer of Ponytail, a fast-growing repository of instruction files for coding agents, has corrected a benchmark after a contributor challenged the original test, InfoQ reports. The revised agentic run reduced the headline code-reduction claim to 54%.
Ponytail drew attention by promising to make coding agents stop over-building. Its early claim of 80% to 94% less code was based on a flawed baseline, according to the update. The maintainer then rebuilt the benchmark around a more realistic agent run.
The correction is useful because many AI developer tools spread through benchmarks, screenshots, and viral claims before their methods are closely examined. A lower but better-tested number gives users a clearer basis for judging whether the technique helps.
The episode also shows why reproducible evaluation matters for agent tooling. Instructions can change model behavior, but the measured effect depends heavily on the task, baseline, and evaluation setup.