Francis deSouza published a number this week that deserves more than a LinkedIn like. Scale AI’s CliniCARE-Bench, built on 750 real patient cases and 25 clinical care scenarios, found that the underlying models in agentic clinical systems reached error rates of 34.7%. When Scale required those same systems to be both correct and free of incorrect shortcuts, scores dropped by as much as 14.8 percentage points. The agents were arriving at the right answer without reading the chart. They were skipping the longitudinal record, bypassing conflicting evidence, and presenting a conclusion that looked earned but wasn’t. In a benchmark environment, that’s a finding. In a live trial, it’s a protocol deviation no one detected.

That gap, between a system that scores well and a system a clinician can operationally rely on, is precisely what the clinical trials field has not solved. And right now, sponsors, CROs, and health systems are deploying agentic AI anyway.

The Governance Deficit Is Structural

The benchmark problem would be containable if governance frameworks were keeping pace. They aren’t. Mo Johnson, a cardiothoracic surgeon tracking AI accountability, cited a February survey of 120 health systems that should stop every AI deployment committee in its tracks: 75% have deployed AI or plan to, 18% have mature governance with a documented strategy and a formal enforcement group, and 42% have neither. Not a lean framework. Nothing structural underneath them at all.

Johnson draws a distinction the industry has been treating as semantic when it is actually operational. Healthcare AI, the administrative layer managing scheduling, billing, and bed flow, carries operational failure costs. Clinical AI, the decision layer covering diagnostic imaging, treatment planning, sepsis prediction, and trial endpoint adjudication, carries patient outcome failure costs. Most governance frameworks treat them identically. That means a health system can deploy an agentic tool that touches protocol eligibility criteria or safety signal detection under the same oversight structure it uses for a revenue cycle automation. The accountability stakes are not remotely comparable.

Cassandra Chuljian, writing from the regulated industry perspective, made the failure mode explicit: “the scariest AI agents aren’t necessarily the ones that fail. They’re the ones that seem to work.” An agent producing reasonable-looking outputs while running on stale data or flawed logic won’t trigger an alert. It won’t generate a deviation report. It will produce a clean-looking record that satisfies an auditor until a patient outcome forces the retrospective audit that uncovers the process failure buried underneath the correct-looking answer. That is silent failure in a regulated environment, and it is the specific failure mode that no current FDA guidance on AI-enabled devices, including the agency’s 2021 action plan for AI and machine learning-based software as a medical device, has operationally resolved for agentic systems operating inside trial workflows.

Process Fidelity Over Benchmark Performance

Sachin Bajpai published research in the International Journal of Computer this week arguing that the architecture conversation needs to happen before deployment at scale, not during it. His framework, validated against a sepsis-deterioration scenario, separates the thinking layer from the acting layer, matches human oversight to the level of decision risk, and builds the audit trail into the architecture from the start rather than retrofitting it as a compliance checkbox. The third element is the one most clinical operations teams skip. An audit trail added post-deployment is a documentation artifact. An audit trail built into the agent’s decision architecture is an accountability mechanism. Those are not the same thing.

Sumant Ranji, Director of UCSF’s Coordinating Center for Diagnostic Excellence, framed the deeper epistemological problem by referencing a randomized trial of an LLM-based decision-support system in Kenyan primary care clinics, published in Nature Medicine. The AI improved documentation of diagnoses and treatment plans, a measurable process outcome, but did not affect the primary clinical outcome of treatment failure. The process improved. The outcome did not move. Ranji connects this to Donabedian’s quality triad and to decades of quality improvement work demonstrating that measurable process gains do not automatically translate into outcome gains. The clinical AI field is at risk of repeating exactly that lesson, this time at the speed of agentic deployment.

This is the counterintuitive claim the industry needs to sit with. Better benchmark performance may actually increase deployment confidence in systems that fail on the dimensions that matter most for patients. A system that scores well on CliniCARE-Bench while bypassing the longitudinal record 14.8% of the time will be approved for broader use faster than a system that scores lower but documents its reasoning at every decision node. The benchmark rewards the outcome. The trial rewards the process. Scale AI is to be credited for building a benchmark that surfaces this tension explicitly, but surfacing it and solving for it in deployment governance are different problems entirely.

Who Answers When the Agent Is Wrong

Rubén Lozano Aguilera, writing from Ai2, made the accountability argument most precisely: “AI is not a moral agent; it can be reliable or unreliable, but it cannot be trustworthy.” Calling an agentic clinical system trustworthy, as vendors routinely do in their deployment materials, displaces the accountability question onto the technology and away from the humans who built it, validated it, and authorized its use in a trial workflow. The Ai2 partnership with Providence Swedish and the Earle A. Chiles Research Institute in Portland offers a more honest model: Asta AutoDiscovery flagged that invasive lobular carcinoma, a breast cancer type historically excluded from immunotherapy trials as immunologically cold, showed more immune activity than previously characterized in the TCGA dataset. The scientists were skeptical, per Lozano Aguilera’s account. They confirmed the signal in an independent patient dataset and in real tumor tissue before committing to a clinical trial. The AI located the signal. The humans defended the inference. That division of accountability is not a limitation of the system; it is the design.

Ingrid O’Dwyer at Sanofi noted this week that approximately 117 AI-discovered drugs have now entered clinical trials, with Insilico Medicine’s rentosertib as the most prominent example. Per O’Dwyer’s read of the data, AI drugs are not outperforming traditional molecules on clinical success rates, but several startups have compressed discovery-to-Phase 1 timelines from the typical 4 to 4.5 years down to 1.5 to 2 years. Speed is real. Superiority has not arrived. That distinction matters because the governance frameworks being built now, while AI drugs are fast but not demonstrably better, will be the same frameworks in place when and if the performance gap closes and the deployment stakes increase further.

Michelle Longmire at Medable pointed to Tufts Center for the Study of Drug Development data showing ROI from agentic AI in oncology trials, describing the moment as a watershed. It may be. But watersheds are also the moment when downstream infrastructure either holds the load or reveals it was never designed for it. The accountability infrastructure for agentic AI in clinical trials, who owns the failure, what the audit trail must contain, how silent process errors get detected before they compound, has not been designed for the volume of deployment the field is now executing. CliniCARE-Bench gave the industry a precise vocabulary for the failure modes. The 42% of health systems running clinical AI with no governance structure underneath it suggests the vocabulary hasn’t reached the room where deployment decisions are made.

The next regulatory action on AI-enabled clinical tools will not cite benchmark scores. It will cite a patient outcome, a missing audit trail, and a sponsor who could not explain how the agent reached its conclusion. Build the architecture that can answer that question before the question gets asked in a Form 483.

References

  1. Francis deSouza, LinkedIn post on CliniCARE-Bench, Scale AI Labs
  2. Mo Johnson, LinkedIn post on clinical vs. healthcare AI governance, Eliciting Insights survey of 120 health systems
  3. Sachin Bajpai, LinkedIn post on agentic AI architecture, International Journal of Computer
  4. Cassandra Chuljian, LinkedIn post on silent AI failure modes in regulated industries
  5. Sumant Ranji, LinkedIn post on AI quality measurement and Nature Medicine LLM trial in Kenyan primary care
  6. Rubén Lozano Aguilera, LinkedIn post on AI trustworthiness and Ai2/Providence Swedish partnership, Asta AutoDiscovery and TCGA dataset
  7. Ingrid O’Dwyer, LinkedIn post on AI drug performance in clinical trials, Insilico Medicine rentosertib
  8. Michelle Longmire, LinkedIn post on agentic AI in oncology trials, Tufts Center for the Study of Drug Development
Website |  + posts

Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.