Frontier web agents fully completed fewer than 3% of the long, open-ended tasks in a new benchmark called KNOWS. Unlike tests that stop after finding a fact or clicking the correct control, KNOWS asks an agent to research across several steps and turn the findings into a coherent document, presentation, or spreadsheet.

Each task combines information retrieval, synthesis, planning, and visual interaction with software. Evaluators mix deterministic checks with language-model judgments so they can assess both concrete requirements and the quality of a final artifact. The tested agents often earned moderate partial-credit scores, showing they could complete many individual steps.

Partial completion was not enough to produce useful work. A single visual failure could make the resulting artifact unusable even when an agent passed more than half of the other checks. As with any new benchmark, the task selection and automated evaluators may not mirror every office workflow. The result nevertheless highlights a gap hidden by short computer-use tests: retrieving information is only the beginning. A dependable assistant must organize evidence, operate unfamiliar interfaces, recover from mistakes, and deliver a file that another person can actually use.