A month-long early-access test found GPT-6 Astra capable of handling much of an AI engineer’s operational workflow, according to Latent Space. The team says it used more than 20 billion tokens while asking the model to choose and train models, label data, monitor pipelines, inspect logs, deploy systems and coordinate other agents.

The most notable change was not a single benchmark result but the model’s ability to keep long projects coherent while managing specialized subagents. Testers report using it to build internal tools, replace several paid software services and run experiments involving model training and evaluation.

Latent Space estimated that one continuously running agent cost less than $6 an hour at preview pricing and observed output speed. That figure is not a ceiling: the model often launched 20 to 50 agents in parallel, which increased spending substantially. The analysis also warns that preview latency may differ when access becomes generally available.

These are hands-on observations from one unusually large test, not an independent production benchmark. Still, they suggest that frontier coding models are expanding from writing code toward supervising complete AI-development loops. Teams considering the model will need to judge total workflow cost, reliability and oversight rather than comparing token prices alone.