Akka has tested a specification-driven coding-agent workflow across 65 open-source projects, measuring time, token use, test parity, code size and runtime performance. The initial tranche consumed 9.41 billion tokens over 99.3 hours, and Akka reported a code-size or performance improvement in 57 ports.

The process moved through discovery, specification, porting, benchmarking and revision. All 65 projects were analyzed and up to 10% of each was implemented before 10 were selected for complete ports. Validation reused original unit and integration tests and added checks for serialization, security, errors, personal data, idempotency and architectural boundaries.

Model behavior differed. Sonnet averaged 61 minutes per port versus 120 minutes for Opus, while Opus used about 40% fewer tokens. Raising the reasoning-effort setting increased consumption without consistently improving efficiency. Akka found that structured claims, evidence and typed behavior improved first-pass work, while missing cross-component context remained a problem.

The headline improvements require caution. Infrastructure and tooling projects showed median performance degradation, Netflix Metaflow was about 100 times slower, and Akka says the workloads behind one reported 143,333-fold gain were different. The experiment supports rigorous specifications and validation, not a blanket claim that automated ports improve software.