METR has introduced a metric called the expenditure horizon, The Decoder reports. The measure tries to put a dollar figure on when AI agents become more expensive than humans for solving a given kind of problem.
That framing is useful because agent evaluations often emphasize whether a model can complete a task, while companies also need to know whether it is economical. An agent that succeeds only after many retries, long runtimes, or expensive model calls may be less attractive than a human workflow.
The Decoder notes that early results on a NanoGPT speedrun are underwhelming and that the metric has blind spots. The idea is still valuable: as agents move into production, cost-effectiveness may become as important as raw capability scores in deciding where automation actually makes sense.