Google DeepMind expands Co-Scientist to run lab equipment
Google DeepMind has upgraded its Co-Scientist AI into a closed-loop research partner that can run physical lab equipment and write papers, marking a major step toward automated science.

Google DeepMind has expanded its Co-Scientist multi-agent system, transitioning it from a simple hypothesis generator into an active laboratory partner. Built on Gemini models, the system now manages a closed-loop workflow that designs experiments, controls physical lab hardware, and writes scientific papers. In materials science, researchers paired Co-Scientist with a high-temperature furnace. Using Gemini 3 Deep Think for direct equipment control, the system cut recipe development time from days to minutes, successfully synthesizing three semiconductor thin films on its first attempt. After 25 rounds of human-refined testing, it also mapped a safer production pathway for a coveted 2D material.
The system demonstrated its versatility across other scientific fields. In biology, Co-Scientist utilized Gemini 3 Pro Image to construct an image analysis pipeline predicting the growth patterns of genetically engineered E. coli colonies, matching unpublished laboratory data for three out of four shape features. In a fully autonomous computer science trial, the system designed a medical AI architecture named Agent_H. While Agent_H outperformed six frontier models, including GPT-5 and Claude Opus 5, on automated health benchmarks, human validation revealed limitations. Three board-certified physicians evaluating the system found that Agent_H only significantly outperformed a baseline Gemini 3.1 Pro model in reducing potentially harmful responses, while automated evaluator Gemini 3.5 Flash showed low correlation with human judgments.
To combat the high hallucination rates typical of autonomous AI agents, DeepMind integrated verification modules that cross-check text claims against code execution logs. In a double-blind evaluation involving 30 domain experts and 450 reviews of 150 AI-generated papers, Co-Scientist fabricated key results in just 4 percent of cases when these modules were active, compared to 46 percent without them and 90 percent in a baseline system. Near-plagiarized content fell from 60 percent to 16 percent, and its safety architecture blocked 98.7 percent of dangerous research paths. For practitioners, this development signals a shift from AI as a mere writing assistant to a collaborative lab agent, though researchers must still watch for selective reporting and discrepancies between generated text and actual code.
This is our own summary of reporting by The Decoder



