Open the FDA’s device authorization database and search for AI-powered clinical decision support tools. Then try to find the published clinical validation data for each one. A multi-institutional study from UNC, Duke, Oxford, Columbia, and the University of Miami did exactly that, and what they found should stop any clinical operations leader cold: 43% of 521 FDA-authorized AI-driven clinical devices had no publicly available clinical validation data. These tools are already inside hospitals. They are already influencing care decisions. And nearly half of them have never been independently tested in a published study.
That is the signal. The research network that Emily Tat of Columbia University Irving Medical Center and NewYork-Presbyterian Hospital and Peter Brodeur of Beth Israel Deaconess Medical Center are building, profiled in a JAMA+ AI Conversations piece, represents one of the most structurally serious attempts to address that gap from the inside out. Their work focuses on the evaluation of clinical AI: not building it, but stress-testing it in ways the FDA currently cannot require and the market currently does not reward.
The Validation Gap the FDA Left Open
The old assumption was that FDA clearance functioned as a quality signal. If the agency authorized a tool, it had been validated. Sponsors, health systems, and trial operators accepted this logic because the alternative, running independent evaluation studies on every deployed AI system, seemed operationally impossible at scale. That assumption is now untenable.
The FDA’s AI/ML-Based Software as a Medical Device Action Plan, released January 12, 2021, laid out a framework for oversight of adaptive AI systems. It was a serious document, and it moved the conversation forward. But it did not mandate post-market clinical validation studies, and it did not create an independent evaluation infrastructure. Five years later, the FDA’s 2024 white paper on AI across CBER, CDER, CDRH, and OCP extended the coordination posture but still stops short of requiring the kind of prospective, site-specific performance testing that would tell a clinical site whether a tool actually works in their patient population.
The market has not filled that gap either. The AI-powered clinical decision support market was valued at $730 million in 2024 and is projected to reach $1.79 billion by 2030, growing at a 15.6% CAGR. At that growth rate, the number of unvalidated tools in clinical environments will compound faster than any regulatory body can audit them retroactively.
This is the structural problem the Columbia-Beth Israel network is trying to solve before the market outruns the governance entirely.
Who Gets Exposed First
Clinical trial operations sit at the sharpest edge of this risk. When a clinical AI tool influences site-level decisions, such as flagging protocol eligibility, predicting dropout, or supporting endpoint adjudication, it becomes part of the evidentiary chain that the FDA will scrutinize during a Bioresearch Monitoring (BIMO) inspection. An unvalidated tool embedded in that chain is a GCP compliance exposure with a dollar sign attached to it.
Therapeutic areas with high protocol complexity face the highest near-term risk. Oncology trials using AI-assisted imaging reads, CNS trials relying on digital biomarkers for endpoint capture, and cardiovascular studies using algorithmic safety signal detection are all operating with tools that, in many cases, have cleared the FDA without a single published independent validation study. The April 2, 2026 FDA Warning Letter 320-26-58, the agency’s first to explicitly cite inappropriate AI use in pharmaceutical manufacturing, signals that the enforcement posture is shifting. What began in manufacturing will reach clinical operations.
Decentralized trial designs face a compounding version of this problem. DCT platforms are increasingly built on AI layers for remote monitoring, ePRO anomaly detection, and site-less eligibility screening. When those AI components lack external validation data, the entire DCT architecture carries an embedded assumption of trustworthiness that has not been tested against the patient populations it is actually serving. CMS recognized a version of this risk when, as of February 6, 2024, it clarified that Medicare Advantage plans using AI for coverage determinations must still meet individualized care standards, a signal that algorithmic outputs alone will not satisfy reviewers.
Sponsors running adaptive trials are in a particularly exposed position. If an AI tool influences adaptive decision rules during an ongoing trial, and that tool’s performance in the specific enrolled population was never independently benchmarked, the adaptive framework rests on a foundation that cannot be audited. The FDA’s statistical reviewers will notice.
The Operational Directive
If you are currently deploying any clinical AI tool for endpoint adjudication, eligibility screening, safety monitoring, or data quality flagging, pull the vendor’s validation documentation now and ask one question: was this tool validated in a patient population comparable to your trial’s enrolled population, and is that validation published and independently reproducible? If the answer to either half is no, you have a protocol risk that belongs in your risk management plan, not buried in a vendor agreement. The NSF-NIH Smart Health and Biomedical Research program is actively funding evaluation infrastructure, but the grants cycle faster than your next BIMO inspection will arrive.
The work Tat and Brodeur are doing through their research network represents the kind of independent evaluation architecture that sponsors should be actively engaging with, not waiting for regulators to require. Submitting your deployed tools to an external evaluation consortium before the FDA asks you to defend them is the kind of proactive compliance posture that protects both the trial and the patients in it.
Watch for the FDA’s next iteration of its AI/ML SaMD guidance, expected to address post-market performance monitoring requirements with more specificity than the 2021 Action Plan. When it drops, any sponsor who has not already established an independent validation baseline for their clinical AI stack will be scrambling to build one on a deadline. The question clinical ops leaders should be asking right now is whether they want that baseline to exist because they built it, or because the FDA told them they had to.
References
- JAMA+ AI — “Designing Trustworthy Clinical AI” (JAMA+ AI Conversations, 2026)
- Medical Economics — “Nearly Half of FDA-Authorized AI Tools May Be Ineffective” (citing multi-institutional study in Nature Medicine)
- FDA — “Artificial Intelligence/Machine Learning (AI/ML)-Based Software as a Medical Device (SaMD) Action Plan,” January 12, 2021
- FDA / King & Spalding — “Artificial Intelligence & Medical Products: How CBER, CDER, CDRH, and OCP Are Working Together,” 2024
- Mordor Intelligence — “AI-Powered Clinical Decision Support Market Forecasts to 2030,” June 13, 2025
- Manatt Health — “CMS Weighs In on the Use of Algorithms and Artificial Intelligence,” February 2024
- TeleDirect MD — “FDA First AI Warning Letter 2026,” citing Warning Letter 320-26-58, April 2, 2026
- NSF — “Smart Health and Biomedical Research in the Era of Artificial Intelligence and Advanced Data Science (SCH),” Solicitation NSF 25-542
Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.

