GPT-6 Astra has become the first model to beat Andon Labs’ human-and-AI reference code on every part of Drone-Bench. The test asks models to write software that maps an office, locates a drone, navigates, identifies a specified person and follows that person using a low-cost DJI Tello EDU.

The achievement came from Astra’s best submissions, not one reliable end-to-end attempt. It beat the reference on person detection in four of ten runs and on three-dimensional reconstruction in one of ten. Andon estimates that an average run has only a 2.8% chance of passing all five stages in sequence. The lab operates the benchmark itself and does not give model developers access to it.

Astra also led six runs of Vending-Bench, a simulated year-long business task. Starting with $500, it averaged a final balance of $15,515, versus $5,422 for Claude Fable 5.1. Andon reported that Astra negotiated more consistently, avoided identified prepayment losses to closed suppliers and refused a price-fixing proposal in three competitive games.

Those results measure behavior in controlled scenarios, not safe or dependable operation in the real world. The drone test is particularly sensitive because its tasks include locating and following a person. Best-case benchmark progress shows what code the model can produce, while the low combined success rate shows how far it remains from routine autonomous deployment.