ZICQ
中 Log in / Sign up
Newsroom Industry & Trends #IA DRUG DISCOVERY #Data quality #Autonomous laboratory #Cytiva #Eroom Law

AI Drug Discovery Hits Data Wall: Cytiva Expert Calls for Closing the Data Loop

Avatar of Mr.Xu

By Mr.Xu

Published:

中文阅读 (Chinese) English Version

Summary:MIT Technology Review Insights publishes an in-depth industry analysis focusing on the current state and challenges of AI in drug discovery. The article notes that while AI significantly improves efficiency in hit identification, model training hits a 'data wall' due to lack of negative data, publication bias, and increased risk of data fabrication with generative AI. Paul Belcher, Director of Protein Research Strategy at Cytiva, emphasizes that high-quality, complete data and lab system integra


Closing the Data Loop in AI-Driven Drug Discovery

Drug discovery is a high-cost, high-risk endeavor. Since the 1950s, the cost of developing new pharmaceuticals has roughly doubled every nine years (Eroom's Law). Today, bringing a new drug to market takes 10-15 years and costs $1-2.5 billion, with failure rates above 90%. AI is the industry's biggest bet to improve success rates and compress timelines.

AI Brings Efficiency to the Lab

One of the most promising early-stage applications is hit identification. Instead of physically screening libraries, AI can design drug candidates from scratch and predict interactions with disease targets, eliminating low-quality candidates before physical testing. Paul Belcher, Director of Protein Research Strategy at Cytiva, notes that AI increases the number and quality of hits, but also demands higher-throughput, information-rich technologies to validate and characterize these complex AI-generated compounds.

Models Need Complete, Quality Data

As AI accelerates demand for data-rich lab systems, a fundamental need emerges: better, more complete data. Many early AI models trained on public datasets are hitting a 'data wall'—all models access the same data, leading to similar conclusions and diminishing returns. Public datasets lack structure, labeling, and diversity, and are not designed for AI.

Publication bias reinforces the problem: most datasets and publications focus on positive results, while negative data (failed experiments, compounds that don't bind) remains buried in lab notebooks. Belcher jokes about a 'journal of negative data.' Without negative data, models cannot adequately learn to avoid bias.

Generative AI has made data fabrication easier. For instance, Western blots are common targets for manipulation; Elisabeth Bik found almost 4% of biomedical papers contained duplicated or manipulated images in 2016. Manipulated data could have disastrous consequences when used to train models. Vendors like Cytiva offer solutions such as Image Integrity Checker, which uses secure hash algorithms (blockchain technology) to detect tampered images, gaining interest from publishers.

Autonomous Labs Could Accelerate Breakthroughs

Belcher envisions fully autonomous 'dark labs' that run with minimal human intervention, cycling through prediction, testing, and optimization, feeding results back into AI models. This could improve success rates of drug candidates entering clinical trials. But automation depends heavily on integration: interoperable systems, structured datasets, and smooth information flow. Most labs are not there yet; many instruments are standalone closed ecosystems.

Integrated infrastructure can enable labs to generate FAIR (findable, accessible, interoperable, reusable) data at scale, training subsequent AI models and closing the loop between dry and wet labs.

On Costs and What Comes Next

AI-driven drug discovery is still early; no AI-discovered drug has yet received full FDA approval, but Belcher expects that to change in 2-3 years. The holy grail is full in silico prediction of efficacy and toxicity, eliminating most wet lab work, but barriers include model maturity, regulatory hurdles, and costs. Stanford study shows training costs for frontier AI models have more than doubled every year since 2016, adding financial pressure.

Belcher remains optimistic: 'As long as the cost of compute doesn't outweigh the cost of clinical development, AI is going to be an advantage.'

This content was produced by Insights, the custom content arm of MIT Technology Review. It was not written by the editorial staff.

Source: MIT Technology Review

— END —

Tags: #IA DRUG DISCOVERY #Data quality #Autonomous laboratory #Cytiva #Eroom Law

Community Comments

Loading live comments and annotations…