Google DeepMind has expanded Co-Scientist into a closed-loop research system that can move from a question to hypotheses, experimental plans, code or machine-readable lab instructions, result analysis and a draft paper. Google reports validated demonstrations in materials science, biology and computer science, with different levels of human involvement.

In materials work, the system proposed a safer route to a two-dimensional material and produced equipment-specific recipes, though the target atomic structure is not yet definitively confirmed. In biology, image-based predictions matched unpublished results for three of four tested shape features but did not generalize beyond known conditions. In computer science, it designed a medical-answering architecture that beat six frontier models on automated benchmarks.

That apparent medical advantage largely disappeared when three physicians judged the answers: the new system was significantly better than the Gemini baseline in only one of nine categories, reduced harmful responses. DeepMind also tested safeguards against invented research results. Across 150 generated papers and 450 expert reviews, key-result fabrication fell from 46% without reliability modules to 4% with them, though mismatches between described methods and actual code remained. Co-Scientist can accelerate parts of experimentation, but the study’s strongest lesson is that execution logs, domain experts and real-world validation remain necessary even when automated scores look convincing.