A physician sits down to review a morning clinic note generated by her ambient AI scribe. The patient, a woman in her fifties managing type 2 diabetes, had spent most of the visit talking about her husband’s worsening dementia, her exhaustion as his primary caregiver, and her mounting fear about what comes next. Almost as an aside, she mentioned that her pharmacy had switched metformin brands and she was not sure the new tablets were the same medication, so she stopped taking them. The AI captured the conversation in full. Its note read: “Discussed glycaemic control.”
That is not a transcription error. The AI did not malfunction. It made a judgment about what the clinical record should contain, and it judged wrong.
Across clinical operations circles, that Lancet vignette is being treated as a cautionary anecdote about physician burnout and bedside manner. In regulatory affairs, it should be read as a source documentation crisis hiding in plain sight.
The Panic Everyone Is Having
The backlash against ambient AI scribes has been growing louder for months. Critics warn that automating clinical note generation strips the therapeutic encounter of its humanity, depersonalizes care, and creates legal liability when a physician signs a note they did not write word for word. Some patient advocacy groups have raised concerns about AI listening devices in examination rooms without meaningful informed consent processes. The American Medical Association’s 2025 policy statement on AI in clinical settings called for “robust physician oversight” of AI-generated documentation before finalization, framing the issue primarily as one of accuracy and liability.
These concerns are legitimate. But they are aimed at the wrong target.
The loudest version of the fear, that ambient AI will hallucinate information and fabricate clinical facts, is the version most easily refuted by the data. A two-month pilot conducted from July to August 2024 involving 31 physicians across multiple specialties generated 7,545 clinic notes using an ambient AI scribe; physicians reviewed 356 of them, representing 4.7% of the total. The error rates, while present, were not the catastrophic fabrication events the critics imagine. A cross-sectional evaluation published in the Annals of Internal Medicine on April 17, 2026 found AI-generated notes were lower in quality than those produced by human clinicians across 11 scribe tools tested against audio recordings of five standardized primary care visits, with 18 human clinicians serving as comparators. Lower quality, yes. A systemic generator of phantom diagnoses, no.
The regulatory infrastructure already anticipates this class of risk. The FDA’s October 2024 guidance on electronic systems in clinical investigations, analyzed in detail by Foley in November 2024, requires audit trails, access controls, and validation procedures for any electronic system that captures source data. ICH E6(R3), the current Good Clinical Practice standard, applies the same logic: data, context, and audit trail are inseparable. If ambient AI is deployed as a source data capture tool in a trial setting, those validation requirements attach to every note it produces. Sponsors who have done their due diligence have already begun vendor qualification exercises for scribe platforms. The scaffolding exists.
So the critics are worried about hallucination, and the regulators have built a framework for electronic data integrity. Both of them are, in a narrow technical sense, correct to do so.
Gradual, however, does not mean harmless.
The Omission Nobody Is Auditing
The real failure mode is not fabrication. It is systematic, structurally invisible omission of precisely the information that determines whether a trial’s conclusions reflect biological reality or controlled-setting artifacts.
Return to the Lancet patient. She stopped taking metformin because her pharmacy switched brands and she did not trust the substitution. That sentence contains three protocol-relevant data points for any diabetes outcomes trial in which she might be enrolled: an unplanned medication interruption, a patient-reported adherence barrier rooted in health literacy and trust, and a social context (sole caregiver for a cognitively declining spouse) that predicts future non-adherence regardless of what the protocol mandates. The AI rendered all three invisible under the phrase “discussed glycaemic control.” That note would pass an audit. The audit trail would show a physician-signed, AI-generated entry with a timestamp and an access log. ICH E6(R3) would see nothing wrong.
This is the structural gap. Every validation framework governing eClinical documentation asks whether what was recorded is accurate. None of them are designed to detect whether what mattered was recorded at all.
The bias runs in a predictable direction. Research from Mass General Brigham found that finely tuned large language models could identify 93.8% of patients with adverse social determinants of health from clinicians’ notes when those notes contained the relevant language. The problem is that official diagnostic codes included those same social factors in fewer than 20% of cases. The AI is only as complete as the documentation it trains on and, in the ambient scribe context, only as complete as what its language model treats as clinically salient. A model trained on structured EHR notes from a health system where psychosocial factors are chronically under-coded will systematically under-generate those elements in its output. The omission is not random noise. It is directional.
For sponsors running trials with patient-reported outcome endpoints, adherence-sensitive pharmacology, or diversity-focused enrollment targets, this directionality is a validity threat. A Phase 3 hypertension trial that enrolls caregiving-burdened patients whose stress load is never captured in source documentation will generate adherence data that looks cleaner than it is. The treatment effect estimate will be calculated on a population that looks, on paper, like it received the intervention as prescribed. The real-world replication of that effect after approval will be considerably messier, and no one will be able to trace the discrepancy back to what an AI scribe chose not to write in 2026.
The EMA has stated explicitly in its guideline on computerised systems and electronic data in clinical trials that “data, contextual information, and the audit trail should not be separated.” Contextual information. The EMA anticipates that context travels with data. Ambient AI, optimized for clinical efficiency rather than regulatory completeness, is severing that connection one encounter at a time, and the severance is not appearing in any deviation log.
Where the Framework Breaks Down
The ICH E6(R3) GCP framework and the FDA’s October 2024 electronic systems guidance together represent the most rigorous data integrity standards ever applied to clinical trial documentation. They validate systems. They require audit trails. They mandate that source data be attributable, legible, contemporaneous, original, and accurate, the ALCOA standard that every data manager knows by heart. What ALCOA does not contain is a criterion for completeness of clinical context. Attributable, yes. Accurate, yes. But accurate relative to what was said, or accurate relative to what the AI decided to include?
Sponsors deploying ambient AI in trial sites right now are, in most cases, treating these tools as productivity infrastructure rather than source data systems. That framing lets them avoid vendor qualification under 21 CFR Part 11 and sidestep the validation burden the FDA’s 2024 guidance would otherwise impose. But if the AI-generated note is what the investigator signs, and the investigator’s signature attests to source data accuracy, then the note is the source document, and the tool that generated it is a source data system whether the sponsor has classified it that way or not.
The FDA has not yet issued guidance specific to ambient AI in clinical trial settings. When it does, the question of what “accuracy” means for a note that is technically truthful but structurally incomplete will be the hardest thing to answer.
The Lancet physician reviewed her AI’s note about the metformin patient and presumably corrected it. Most physicians, under productivity pressure at a site processing forty patients a day, will not catch what is not there. The audit will show a signed, validated note. The trial database will show a compliant subject. The submission package will show clean source data. And somewhere in the noise of a Phase 3 efficacy analysis, a real adherence signal will quietly disappear into a three-word phrase about glycaemic control.
References
- The Lancet — “The art of fidelity: clinical documentation and ambient AI”
- JMIR Medical Informatics — Pilot Study on AI Scribe Error Rates in 7,545 Clinic Notes (July–August 2024)
- University of Washington Medicine / Annals of Internal Medicine — “AI Scribes Lower Quality Than Human Clinicians” (April 17, 2026)
- Foley & Lardner — “FDA Clinical Investigations Guidance: Electronic Systems” (November 2024)
- Mass General Brigham — “Generative AI Models Effectively Highlight Social Determinants of Health in Doctors’ Notes”
- European Medicines Agency — “Guideline on Computerised Systems and Electronic Data in Clinical Trials”
Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.

