A new benchmark called ORCA measures whether language models can translate code between data-science libraries without changing what it does. The suite goes beyond code generation by requiring functional equivalence across tools and ecosystems.

ORCA-MAIN contains 1,600 curated tasks covering data queries, data manipulation and deep learning. ORCA-PROJECT adds 200 complete-project translations spanning seven task types. Each task includes a reference translation and tests, and the authors used a multistage process to check correctness and test strength.

Even the strongest reported model, Claude Opus 4.6, succeeded on only 56.92% of ORCA-MAIN and 33.67% of ORCA-PROJECT. Translation was easier when source code expressed its operations explicitly. A method that first inferred the source program’s intent raised average success by 4.80 percentage points on focused tasks and 5.33 points on projects. The benchmark suggests developers should run translated code against comprehensive tests rather than treating a plausible rewrite as interoperable.