Computer-use agents that look successful on screen can still leave an enterprise database in the wrong state. ERPBench, a new benchmark for agents operating business software, checks completed tasks against the underlying records rather than trusting visible clicks or confirmation messages.
The benchmark runs on a live, reproducible enterprise resource planning system covering workflows such as finance, procurement, inventory and customer operations. Agents receive screenshots and act through simulated mouse and keyboard input. The accompanying harness can require human approval before consequential actions, although benchmark runs proceed autonomously for consistent measurement.
Tests of six open and closed models found that strong performance on general desktop tasks did not translate into reliable enterprise work. In the starkest result, some agents reached and saved the correct form in as many as 85% of runs but entered the right stored value in as few as 3%. That distinction matters because an incorrect business record may persist without producing an obvious on-screen error.
ERPBench is one benchmark on one reproducible system, not a complete measure of production readiness. It nevertheless demonstrates why evaluating agents by interface progress alone is inadequate. Organizations need state-level checks, audit trails and approval gates before software can safely modify financial, inventory or customer records.