SynHy Article

Generative Medical AI Pilots Need a Live Evidence File

When medical AI reaches real patients before full authorization, a live evidence file should track intended use, patient risk, performance, monitoring, and stop rules.

The Problem Is Evidence Arriving After Use Begins

Generative AI medical tools can change quickly, learn from broader workflows, and produce outputs that affect care decisions. That makes the evidence problem different from a static device that behaves the same way every time it is used.

STAT reported on September 3, 2026 that the FDA's TEMPO pilot is allowing selected generative AI medical devices to reach patients before full marketing authorization. The FDA's own TEMPO announcement says participating manufacturers may request enforcement discretion for certain requirements while collecting and reporting real-world performance data.

The business and clinical answer is a live evidence file. It records what the AI is allowed to do, what patient risk it creates, what data is being collected, and what would stop or narrow use.

Why The Regulatory Problem Is Different

Traditional medical-device review often assumes a defined product, defined claim, defined patient population, and defined evidence package. Generative systems can be less stable because prompts, retrieval sources, model versions, clinical workflows, and user behavior can change performance.

TEMPO is designed around real-world use in chronic disease contexts, including cardio-kidney-metabolic, musculoskeletal, and behavioral health conditions. That is exactly where digital health tools may produce useful daily signals, but it is also where context, adherence, clinical escalation, and patient understanding matter.

A live evidence file helps bridge that gap. It does not replace regulatory authorization, clinical judgment, or professional responsibility. It makes the pilot's claims, limits, and evidence visible while the system is in use.

The Cost Of Weak Evidence

Weak evidence can harm patients, delay regulatory learning, and damage a company's credibility. If a tool reaches patients before full authorization, unclear monitoring can leave sponsors and clinicians unable to distinguish safe underperformance from unacceptable risk.

The costs include patient harm, clinician workload, emergency correction, payer skepticism, legal exposure, withdrawn access, and a slower path for future devices. The organization may also lose the chance to learn from the pilot because the data was not collected in a usable structure.

The evidence file reduces that failure mode by answering basic questions every week. Who used the tool, for what intended use, under which version, with what performance signal, and under which escalation path?

How To Diagnose Pilot Readiness

Start with intended use. The file should state the patient group, clinical condition, user type, care setting, decision supported, output type, and action that should not be taken from the AI alone.

Then map risks. Include false reassurance, false alarm, delayed escalation, biased performance across patient groups, misunderstood output, missed context, excessive clinician burden, privacy exposure, and drift after model or workflow changes.

Finally, identify the evidence source for each risk. Real-world data should not be a vague promise. It should name measures, collection frequency, review owner, privacy boundary, quality checks, and the decision rule for continuing, narrowing, or stopping use.

Options For Sponsors And Providers

The conservative option is a simulated or clinician-shadow pilot. The AI runs alongside care without influencing patient action until performance and workflow burden are better understood.

The middle option is supervised patient use. The AI supports a narrow intended use, clinicians review outputs, and the sponsor reports real-world performance under a defined protocol.

The highest-risk option is broader patient-facing use before the evidence pathway is mature. That may be appropriate only when patient risk is low, the benefit case is strong, the monitoring plan is concrete, and the stop rule is enforceable.

Build The Live Evidence File

The live evidence file should contain the intended-use statement, patient population, excluded patients, model version, retrieval sources, user workflow, training given to clinicians and patients, real-world performance measures, adverse-event route, privacy control, and stop rule.

It should also preserve version evidence. If a model, prompt, clinical protocol, escalation threshold, user interface, or data source changes during the pilot, the file should show the change date and why the prior evidence still applies or no longer applies.

A useful file is operational, not ceremonial. The review owner should be able to open it during a weekly pilot review and decide whether to continue unchanged, adjust the workflow, pause a subgroup, or stop the pilot.

A Worked Example

Suppose a generative AI tool supports remote chronic-condition coaching for patients with low-acuity risk. The tool summarizes patient-reported symptoms, suggests education, and flags cases for clinician review.

The evidence file states that the tool cannot diagnose, change medication, or replace clinician escalation. It tracks false negatives, false positives, patient comprehension, escalation timeliness, clinician override rate, unresolved safety events, and performance by relevant patient groups.

If the tool misses two predefined escalation triggers in a week, the pilot pauses that workflow until the sponsor reviews transcripts, updates the prompt or escalation logic, and documents the retest. That is a stop rule, not a dashboard note.

Measures That Prove It Works

Track patient safety events, escalation accuracy, clinician override rate, unanswered high-risk messages, user comprehension, time to review, false-alarm burden, dropout rate, and performance differences across patient groups.

Track evidence completeness too. Each patient-affecting output should be traceable to a model version, approved use, input context, review status, and escalation path. Missing traceability weakens both clinical learning and regulatory confidence.

The most useful measure is decision utility. If the pilot cannot explain what evidence would make the tool narrower, broader, paused, or ended, the evidence file is not yet strong enough.

The Next Step This Week

For any medical or wellness AI pilot, write the intended-use sentence in one paragraph. Include who uses it, which patients it affects, what decision it supports, and what it is not allowed to decide.

Create the first evidence table with five rows: safety signal, performance signal, workflow burden, equity signal, and stop rule. Assign an owner and review frequency to each row.

Then run one tabletop scenario. Ask what happens when the AI gives a plausible but unsafe answer, when a patient misunderstands the output, and when a model update changes performance.

Sources And Method

This article uses STAT's September 2026 report on generative AI devices in FDA's TEMPO pilot, the FDA's TEMPO press announcement, FDA TAP material, and NIST's AI Risk Management Framework.

The analysis focuses on evidence operations for sponsors, providers, and business leaders. It is not medical, regulatory, or legal advice, and it does not evaluate any named TEMPO participant's product.

Source links: STAT, FDA TEMPO announcement, FDA TAP program, and NIST AI RMF.