Picture a conference room in Basel or Cambridge or South San Francisco in early 2025. A chief scientific officer clicks to slide three of her investor deck: a timeline graphic showing AI compressing the traditional 12-to-15-year drug development cycle into something that looks, on paper, like four years. The room nods. The check gets written. This scene played out with enough frequency last year that AI-ML drug discovery companies collectively raised $8.9 billion across 264 financing rounds in 2024 alone, with biotechnology AI attracting $5.6 billion of that total. The momentum feels real because, in several narrow respects, it is real.
Insilico Medicine identified a novel target for idiopathic pulmonary fibrosis and advanced a candidate into preclinical trials in 18 months at a reported cost of $150,000, a process that typically consumes four to six years and tens of millions of dollars. That is a genuine signal, not marketing. Exscientia built a platform that has demonstrated measurable compression at the hit-to-lead stage. And on January 14, 2026, the FDA and the European Medicines Agency jointly published ten guiding principles for good AI practice in medicine development, covering the full product lifecycle from early research through safety monitoring. Regulators are paying attention. The infrastructure is taking shape.
The consensus narrative, then, is one of arrival. AI has moved from theoretical to operational, and the only remaining question is execution speed.
Yet as of August 2026, not a single AI-discovered drug holds full FDA approval. Projections for the first such approval remain pinned to 2026 or 2027, and those dates have been quietly sliding for three consecutive years.
The Promises That Hit Phase II
The most instructive data point in the entire AI drug discovery story is a topical cream for itchy skin. BenevolentAI, one of the sector’s most capitalized and credentialed platforms, advanced BEN-2293, a Pan-Trk inhibitor for atopic dermatitis, through its AI-assisted pipeline into a Phase IIa trial. On April 5, 2023, the company announced that BEN-2293 had failed to achieve statistically significant improvement on either the EASI or NRS primary endpoints. The algorithm had identified the target. The chemistry was AI-assisted. The trial still failed.
One failure does not indict a technology. But it clarifies the terms of the debate. The efficiency gains that AI demonstrably delivers sit almost entirely in the preclinical space: target identification, molecular generation, ADMET prediction, synthesis planning. The moment a compound enters a human being, the algorithm’s advantage erodes against the irreducible complexity of biology, patient heterogeneity, and endpoint selection. That is not a software problem. It is a translation problem, and it is one the $8.9 billion has largely not been spent solving.
Which raises an uncomfortable question: what exactly are the competing stakeholders in this ecosystem actually trying to build?
The venture community wants platform valuations, and platform valuations require clinical proof-of-concept. This creates pressure to advance candidates into Phase I and II before the algorithms generating those candidates have been validated to the standard that regulators will eventually require. The FDA and EMA’s ten guiding principles from January 2026 are encouraging in their intent, but they are principles, not requirements. They do not specify how a sponsor should document algorithmic reproducibility in a regulatory submission, what constitutes adequate training data disclosure, or how a model’s version history should be maintained across a multi-year IND. The space between principle and requirement is where candidates go to generate ambiguous Phase II readouts that neither prove nor disprove the technology’s core claims.
Academic AI developers occupy a different position entirely. Yoshua Bengio and other leading figures in the machine learning community have been explicit that pharmaceutical data secrecy is throttling the technology’s development. The argument, reported by Science|Business, is that firms are not sharing their proprietary datasets, limiting the training diversity that would make AI models genuinely generalizable across therapeutic areas. The academics want open data consortia. The companies want proprietary moats. Those two positions are not reconcilable, and the current funding environment rewards the moat.
Large pharma partnerships add a third vector of conflicting intent. When a Pfizer or a Sanofi signs a multi-year AI discovery partnership, what they are often purchasing is optionality: the right to observe whether the platform produces anything worth licensing, without committing to the operational transformation that would actually integrate AI outputs into their clinical development infrastructure. The press release announces a partnership. The integration work that would make the AI’s outputs regulatory-ready, traceable, and auditable across the eClinical stack rarely follows at the same pace.
The Validation Gap Nobody Budgeted For
Here is the structural problem, stated plainly. AI drug discovery has been treated as a chemistry and biology challenge. The regulatory and data infrastructure challenge has been treated as someone else’s problem, to be solved later, by someone downstream.
It will not be solved later. The FDA’s ten AI principles published in January 2026 make clear that the agency expects algorithmic transparency, documentation of data provenance, and evidence of model performance across the intended use population. That language maps directly onto the kind of audit trail that eClinical systems, specifically EDC platforms, CTMS environments, and data management pipelines built to CDISC/SDTM standards, are designed to produce. But most AI discovery platforms were not architected with downstream regulatory submission in mind. They were architected to generate candidates faster. The assumption was that the regulatory integration could be bolted on.
It cannot be bolted on.
Consider what a sponsor actually needs to do to submit a regulatory package in which an AI algorithm meaningfully contributed to target selection or compound optimization. The agency will want to understand the training data: its source, its completeness, its known biases. It will want to understand model versioning: which version of the algorithm produced the compound that is now in Phase III, and whether that version is still the version being used to guide the program. It will want to understand how the algorithm’s outputs were validated against wet-lab results before the decision to advance was made. None of this documentation is automatically generated by the discovery platform. All of it requires deliberate data infrastructure decisions made at the point of discovery, not at the point of submission.
The counterintuitive reality of AI drug discovery is this: the platforms that have produced the most impressive preclinical efficiency gains may actually be the hardest to translate into regulatory submissions, because their speed was achieved precisely by not building the documentation architecture that submissions require. Insilico Medicine’s 18-month, $150,000 preclinical run is extraordinary. The question a regulatory affairs director at a sponsor considering licensing that program must ask is how much of that $150,000 process generated audit-ready records, and how much of the cost to generate those records is still outstanding.
What the Next Approval Will Actually Prove
The first FDA approval of an AI-discovered drug, projected for sometime in 2026 or 2027, will be treated as a validation event for the entire sector. Investment bankers will cite it. Platform companies will reference it in every deck. And it will be, in one narrow sense, a genuine milestone.
But the approval will tell us almost nothing about whether AI drug discovery works at scale across therapeutic areas, patient populations, and regulatory jurisdictions. It will tell us that one platform, with one candidate, in one indication, built enough documentation to satisfy one review division’s questions. The most advanced AI-discovered candidate currently in late-stage development is Insilico Medicine’s rentosertib, which received Orphan Drug Designation from the FDA for idiopathic pulmonary fibrosis, a rare disease with a relatively permissive regulatory pathway. Orphan Drug designations reduce the evidentiary bar. A first approval earned under those conditions will be celebrated as proof of concept for a technology whose commercial case requires it to work in large chronic disease indications with standard evidentiary requirements.
That gap between the conditions of the first approval and the conditions of the commercial opportunity is not a minor footnote. It is the central question that $8.9 billion in annual funding has not yet answered.
The sponsors and CROs building clinical development infrastructure today face a choice that the current discourse obscures. Investing in eClinical data architecture that can receive, document, and validate AI-generated inputs is not a technology upgrade. It is a prerequisite for any AI discovery program to produce a regulatory package that survives scrutiny. The platforms generating the candidates will not build that infrastructure for you. The FDA’s January 2026 principles will not build it for you. And the first approval, when it comes, will be claimed by every platform as shared vindication, regardless of who did the hard work of making the data submission-ready.
The question worth watching is not which AI platform produces the first approved drug. It is which sponsor builds the data infrastructure to prove it did.
References
- Nature Reviews Drug Discovery — “Artificial intelligence in drug discovery: what it is, where we stand and the path forward”
- Dealforma — “AI-ML Drug Discovery and Licensing: R&D, M&A, Ventures and IPOs 2025 Review”
- FDA — “Guiding Principles for Good AI Practice in Drug Development” (January 14, 2026)
- Pharma Insight — “BenevolentAI BEN-2293 Phase IIa Failure, April 2023”
- Science|Business — “AI-assisted drug discovery held back by private sector secrecy over datasets”
- PMC — “Insilico Medicine: 18-month preclinical timeline and $150,000 cost for idiopathic pulmonary fibrosis candidate”
- OncoDailyTech — “Insilico Medicine’s Rentosertib: FDA Orphan Drug Designation for Idiopathic Pulmonary Fibrosis”
Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.

