A new arXiv paper introduces PhoneHarness, a benchmark and execution harness for mobile agents. The framework evaluates agents that use a mix of GUI actions, device-side commands, and structured tools rather than only taps and swipes.

That matters because real phone tasks often require more than controlling the screen. Agents may need to use app interfaces, system commands, and tool calls while leaving verifiable evidence that a side effect actually occurred.

PhoneHarness reflects a broader shift in agent evaluation toward realistic workflows, mixed action spaces, and durable verification instead of simple final-screen state checks.