A new study covered by The Decoder challenges claims that autonomous AI research is close to being practical.

Researchers gave agents using Claude Opus 4.8 and GPT-5.6 Sol six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpublished NeurIPS papers then rated the agent-produced results as “Reject.”

The important nuance is that the agents were not useless. According to the report, frontier models could handle much of the research engineering process. The failure was in research judgment, creative problem-solving, and knowing when to abandon an approach that was not working.

That distinction matters for labs and companies planning autonomous research workflows. AI agents may accelerate implementation, literature handling, experiments, and drafting, but the study suggests they still need human direction for the hardest scientific decisions. It also gives a more grounded way to discuss autonomy: not whether a model can complete many steps, but whether it can make the right calls when the path is uncertain.