AI-Driven Drug Discovery Faces Challenge of Data Loop Closure
Drug development costs follow Eroom's Law, doubling every nine years. While AI promises efficiency, lab validation of AI-generated candidates remains a bottleneck. Closing the data loop is key.
Drug discovery is a high-cost, high-risk business facing increasing pressure in a market dominated by first-mover advantage. Since the 1950s, the cost of developing new drugs has doubled approximately every nine years—a phenomenon known as Eroom’s Law. Currently, bringing a new drug to market takes an average of 10 to 15 years, costs between $1 billion and $2.5 billion, and has a failure rate exceeding 90%.
AI has emerged as the pharmaceutical industry’s most promising tool to improve success rates and shorten timelines. By accelerating the identification, testing, and optimization of compounds, AI can reduce the costly risks associated with late-stage development failures. Paul Belcher, Protein Research Strategy Director at global life sciences company Cytiva, stated, “The main costs of drug discovery still lie in the clinical stages. Reducing risks and improving success rates is extremely beneficial. AI is expected to not only save time and compress timelines but also deliver higher-quality candidates to clinical trials.”
Eroom’s Law and AI’s Potential
The rising costs in drug discovery have persisted for over half a century. Eroom’s Law, a play on Moore’s Law reversed, highlights the exponential decline in new drug development efficiency. The contributing factors include stricter regulations, increasingly complex target diseases, and the need for large-scale clinical trials.
AI is seen as the game-changer to reverse this trend. Machine learning models can analyze massive chemical databases to predict promising molecules, uncovering candidates overlooked by traditional high-throughput screening. In fact, several AI-generated drug candidates have recently entered clinical trials, transitioning into the proof-of-concept stage.
AI’s Role in Hit Identification
The most promising early application of AI in drug discovery is hit identification—the process of screening molecular libraries against disease-related targets (such as proteins) to find molecules that bind to them. Successful hits provide researchers with starting points for further testing and optimization, ultimately forming the foundation for viable drugs.
Belcher notes a shift from empirical screening to predictive design. Traditionally, physical libraries were screened, but now AI enables the design of drug candidates from scratch, predicting their interaction with targets before initiating research and development. Companies are no longer limited by the volume of physical screening they can perform. “AI removes that limitation,” Belcher explains. Moreover, low-quality candidates can be filtered out before physical testing, saving time and resources.
However, AI still cannot reliably predict the pharmacokinetics or developability of compounds. Consequently, all AI-generated candidates require laboratory validation.
Challenges and Bottlenecks in Lab Validation
Traditional screening workflows are designed for large-scale hit identification but are not equipped to profile the diverse and complex candidates generated by AI in detail. The burden of testing, characterizing, and refining these compounds falls heavily on laboratory teams.
Belcher explains, “Current hit identification technologies use binary or threshold-based methods to screen hundreds of thousands, sometimes millions, of compounds, generating low-fidelity data (Yes/No responses). While AI increases the number of hits, this results in downstream challenges when generating high-fidelity data for deeper characterization.”
High-fidelity data refers to multi-faceted information such as binding affinity, selectivity, solubility, and metabolic stability of compounds. If AI could leverage this data, subsequent design cycles could achieve even greater accuracy. However, laboratories are struggling to keep up with experiments, leaving the data loop incomplete.
Closing the Data Loop
The true value of AI in drug discovery lies in establishing a rapid “data loop” where predictions and experiments feed into each other. AI models predict candidates, laboratories validate them, and the results are fed back to improve the model’s predictive accuracy. The faster this cycle operates, the more efficiently promising molecules can be identified early.
Yet many pharmaceutical companies face challenges in aligning their lab and AI teams. Experimental data is not fully utilized to refine models due to siloed organizational structures, inconsistent data formats, and missing metadata. Companies like Cytiva are working to bridge this gap by integrating experimental equipment with AI platforms.
Belcher envisions the next step: “Automating the feedback of high-quality lab data into models will enable AI to continuously improve and enhance the reliability of predictions.”
Future Outlook and Market Impact
AI-driven drug discovery remains in its early stages, and significant hurdles must be overcome before AI-generated candidates become approved medications. However, efforts to close the data loop are becoming critical to pharmaceutical companies’ competitiveness. In a market where first-mover advantage is key, shortening development timelines by even a few years can extend exclusive sales periods and greatly impact revenue.
Beyond molecular design, AI can also optimize clinical trial operations, such as patient selection and biomarker discovery. Companies like Cytiva are strengthening ecosystems by offering AI-native experimental platforms.
Nevertheless, widespread adoption of AI in drug discovery faces challenges such as data quality, standardization, regulatory acceptance, and bias issues. Particularly, generative AI’s molecular designs raise concerns about patentability and toxicity prediction limitations.
Editorial Opinion
In the short term, closing the data loop in AI-driven drug discovery will become a vital source of competitive advantage for companies. Those that can integrate AI with laboratory processes will gain first-mover benefits by shortening development timelines and improving success rates. If platform providers like Cytiva can drive standardization, the overall efficiency of the industry could see significant improvement.
Currently, AI is primarily utilized in hit identification, but applications in lead optimization and toxicity prediction are expected to accelerate within the next six months to a year. From a long-term perspective, fully automated data loops could lead to a paradigm shift in drug discovery—transforming it from “experimental science” to “predictive science.” AI could independently generate hypotheses, validate them through automated experimental systems, and update models based on the results. This would fundamentally reshape pharmaceutical business models and research organizations.
However, unresolved issues such as overfitting within closed loops and domain shifts pose challenges. Determining the degree of human oversight and intervention required remains a critical question.
The editorial team believes that while closing the data loop is central to the intrinsic value of AI-driven drug discovery, its implementation will require cross-organizational collaboration and significant investment in standardization.
References
- “Closing the data loop in AI-driven drug discovery”, by MIT Technology Review Insights — MIT Technology Review AI, 2026-07-27T11:40:16.000Z (ARR)
- Source URL: https://www.technologyreview.com/2026/07/27/1139667/closing-the-data-loop-in-ai-driven-drug-discovery/
Frequently Asked Questions
- What is the "data loop" in AI-driven drug discovery?
- It is a cycle where AI-predicted candidate compounds are experimentally validated in laboratories, and the resulting high-fidelity data is fed back into AI models to improve prediction accuracy. Accelerating this loop enables efficient early-stage filtering of promising molecules.
- What is Eroom's Law?
- Eroom's Law describes the phenomenon where the cost of new drug development doubles approximately every nine years. It is the reverse of Moore's Law (exponential improvement in semiconductor performance) and is attributed to factors like stricter regulations and the need for large-scale clinical trials. AI is expected to help counteract this trend.
- What is the biggest challenge in AI-driven drug discovery?
- The bottleneck lies in the laboratory validation of AI-generated candidates. Current screening methods struggle to evaluate the diverse compounds produced by AI, and experimental data is not effectively utilized to improve AI models due to organizational silos and lack of standardization.
Comments