A new benchmark covered by The Decoder suggests that AI systems remain weak at realistic knowledge-work tasks, with the best model fully solving only 3 percent of them. The result is a reminder that real work often requires context, judgment, and multi-step execution beyond standard test sets.

The finding matters because many companies are betting on AI agents to automate office workflows. Low full-solve rates imply that deployment needs human review, narrower task design, or stronger tooling rather than broad autonomy.

The benchmark adds useful friction to the hype around agents by measuring performance in conditions closer to actual work.