On Epoch AI’s MirrorCode benchmark, which tests whether AI models can reconstruct full programs without seeing the original code. The most extreme example cited a model working for 19 days on one task at a cost of $2,600.
The benchmark is useful because ordinary coding tests often miss long-horizon failure modes. Rebuilding a 16,000-line toolkit stresses planning, memory, debugging, and consistency over far longer spans than a typical benchmark prompt.
The results show progress, but also a practical warning: even strong models can become expensive and unreliable when the task demands sustained engineering rather than isolated code generation.