HomeTopicsArtificial Intelligence

Protecting Scientific Intent in AI-Enabled Labs by Not Letting AI Set the...

Forget self-driving labs. For scientific AI to thrive, humans must always be in charge, setting objectives, constraints, risks, and interpreting results. [Maite Mueller/Getty Images]

AI is neither saint nor demon. Nor should it be a replacement for human scientists. As AI takes on greater roles in designing, executing, and analyzing experiments and processes, scientists understand that even the best AI needs human supervision.

The big question is how much oversight is needed and whether—or the extent to which—AI interactions should be documented and reported in regulatory filings.

Le Cong, PhD, associate professor, Stanford University, and co-founder of LabOS and MedOS

Although the lab of the future may be envisioned as a self-driving lab, that’s actually a bad idea, noted Le Cong, PhD, associate professor, Stanford University, and co-founder of LabOS and MedOS. Instead, he and leaders in the AI and biopharmaceutical industries see the future of scientific AI as agentic, with humans in charge.

“We think there is a positive trend toward using AI in a way that’s human, rather than as a self-driving lab,” he said. In that environment, AI is simply a tool—albeit a powerful, adaptive one—that can be managed as long as scientists use the right prompts.

Humans in the lead

“Today, much of the scientific research process remains inaccessible to machines,” Cong and colleagues wrote in a recentpaper. Despite automation and some use of AI, “Scientific discovery remains fragmented.” Specifically, AIs lack the tacit knowledge, evolving experimental context, human observations, and adaptive decision-making inherent in human scientists.

Human involvement is needed, therefore, not just to oversee AI-based activities and check the output, but to ask the right questions and to ensure that analyses make sense in context. Specifically, he describes a scientific setting in which an AI would handle an experiment’s execution, and the scientists would be responsible for:

  • Framing objectives
  • Interpreting results
  • Setting constraints
  • Governing risks

“If AI can interpret everything, then it will start to generate fake stuff, right?” Cong asks. “We’ve seen this when AIs begin guessing in an effort to return results and supply citations that don’t exist. There are certain things that are useful for AI to do in the lab.”

But, as last summer’s sandbox breakouts illustrated, an AI needs firm guidelines as to what it can do, where it can access information, and the degree of autonomy it has in meeting a request.

For example, he recommends adding this phrase to instructions: “Any actions not explicitly stated in the protocol need human approval.” That default to human judgment also should apply to determining the risks associated with certain actions, such as editing a human gene, Cong said. With those guardrails, the paper points out, agentic AIs can freely handle “routine execution and coordination across models, instruments, protocols, and laboratory states.”

Cong equates the scientific use of AI to autonomous vehicles, which, according to the Insurance Institute for Highway Safety, have a 68% lowercrash rateper mile traveled than human drivers in the same environment. Given those statistics, he added, “My thought is to elevate humans to setting destinations. We do not need humans to always execute the driving.”

In a university scientific lab, that equates to staffing a principal investigator and trainees, without much of the hierarchy that exists today. In a corporate environment, the hierarchy flattens to scientists who propose, design, execute, and interpret experiments, and a lab manager. The distinctions between senior and junior scientists blur because much of the hands-on work is automated.

The combination of AI and lab automation is expected to reduce human errors. Cong cited a2016 Nature studyof 1,500 scientists. When asked, “’Can you replicate other people’s experiments, and can you replicate your own after a few months?’ approximately 70% could not replicate others’ experiments, and half could not replicate their own!” Cong said. “AI can improve that.”

Cong equates the scientific use of AI to autonomous vehicles, which, according to the Insurance Institute for Highway Safety, have a 68% lower crash rate per mile traveled than human drivers in the same environment. Given those statistics, “My thought is, elevate humans to setting destinations,” said Cong. “We do not need humans to always execute the driving.” [Chesky_W/Getty Images]

Reasons for such poor reproducibility rest in the details, he elaborated. Was a step omitted? Was the protocol followed exactly? Is the protein being used identical to the one in the original experiment? Were the temperatures the same? Details like this—which AI can duplicate precisely—are behind many reproducibility challenges.

AI risks minimal

As yet, it’s unclear how AI involvement in experiments should be preserved and reported in regulatory submissions, Cong continued. “We’re still early in this journey.”

That said, the risk that AI will escape its constraints and cause physical harm—like designing and developing a physical virus—appears relatively low, according to Cong. That’s because a rogue AI still needs a human accomplice to allow a virus, for example, to be manufactured and released. “In areas where there is a physical execution step, I think AI is still incapable,” he said, “although we are seeing progress in connecting AI to biomedical labs and applications, and the physical execution layer.”

That’s due to the fact that there are multiple layers of human intervention needed to actually manufacture a product. Aside from logistics, he cites good manufacturing practices, safety and efficacy regulations, and digital safeguards like track and trace and the FDA’s 21 CFR Part 11, as well as real-time monitoring, periodic inspections, and quality control activities. Those regulations and checkpoints should also be sufficient to manage variations that occur during manufacturing as real-time conditions drift from specifications.

The catch, as last summer’s breakouts of frontier AIs underscore, is that sometimes AIs exceed their parameters. Whether there is sufficient appreciation of this among AI users remains to be seen, Cong said.

“A lot of people are connecting AI systems, which have access to more and more information and key decision-making systems,” Cong pointed out, without deeply understanding the risks and establishing appropriate guardrails. “People might be overly trusting of AI, perhaps.

“The more powerful the AI, the more capable it is of doing something. Are people keeping pace [with the technology and its risks]?”

Sometimes, small, highly specific AIs may be a better choice than always leveraging the large frontier models, Cong suggested. The reason, Cong, senior corresponding author Mengdi Wang, PhD, professor, Princeton University, and a dozen colleagues, noted in a 2025paperinNature Biomedical Engineering,is that “Large language models often lack domain-specific knowledge and struggle to accurately solve biological design problems.”

Whatever level of AI is used, however, “Humans need to be in the lead throughout the process,” Cong stressed.

NewsArtificial intelligenceAutomationClinical laboratory techniqueGood manufacturing practiceQuality control

Previous article

Built Without Bacteria: Bringing Cell-Free Synthesis to the Bench

Next article

Scientists Analyze 267 Receptors That Control Protein Fate in Rare Diseases

Also of Interest

CloudScope Enables Continuous Remote Monitoring of Brain Activity in Freely Moving MiceSmarter Cell Culture Starts with Better MediaRevvity Signs Agreement to Acquire Human Cell DesignTop 10 Contract Development and Manufacturing Organizations 202610 CDMO Up & Comers 2026Harnessing Real-Time Data

Related Media

Data Integrity as the Foundation for AISmall Molecules, Big Expectations: How CDMOs Are Helping Sponsors Navigate Complexity, Speed, Scale-Up, and SustainabilityBeyond the Technology: AI Readiness in Regulated LaboratoriesAACR 2026 Video Update: Cancer Research Edges Toward an AI-Driven EraRobots on the Red Line: A Video Update from SLAS 2026Reporting Live from JPM 2026: Alex Philippidis and Jonathan Grinstein, PhDTop 5

ResourcesRecommended For You

Podcast

Touching Base

Touching Base is the dynamic podcast series from the editors ofGEN. Each episode features a rotating case of senior editors—including John Sterling, Kevin Davies, Julianna LeMieux, Alex Phillippidis, Uduak Thomas, Corinna Singleman, and Fay Lin—who delve into emerging stories, exchange ideas, and debate the latest trends in biotech. Additionally, they talk to some of the leading voices in the industry about what's now and next.Start listening today!

Stay up to date with the lasted episodes of Touching Base bysubscribing to theGENPodcast Newsletter