Toggle light / dark theme

AI-powered medical devices must be tested in real-world settings

@ Nature this week calls out lack of real world data for AI medical algorithms.

AI-powered medical devices that inform clinical decision-making need rigorous, real-world testing equivalent to what’s required for drugs and self-driving cars.

- Since ChatGPT’s release in late 2022, generative AI has flooded into clinics rapidly — roughly 3 peer-reviewed articles on clinical AI publish every day, and 230+ million people weekly ask ChatGPT health questions.

- AI systems now handle admin tasks, order labs, help prescribe drugs, interpret X-rays/MRI/CT scans, and diagnose rare diseases.

- But regulatory oversight lags: some AI-powered medical products are being certified *without* real-world assessment.

Why current testing falls short:

- Most AI tools are novel, so there’s little existing clinical data to benchmark against (unlike a revised stethoscope or new bandage).

- A systematic review found only 23% of ~4,600 papers on clinical AI used real-world patient data, and just 19 were prospective randomized trials.

- Companies often rely on one-off simulated scenarios comparing AI judgment to physician judgment — and studies comparing specialized vs. general-purpose models show conflicting accuracy results.

- Patient perspectives are largely ignored (patients aren’t asked whether they want these tools in their care).

The regulatory landscape:

- The FDA published a discussion paper in August seeking feedback on regulating generative-AI-enabled medical devices (deadline: 19 October). The editorial urges researchers to submit views.

- Current FDA approach uses a risk scale: low-risk items (bandages) can be self-certified; mid-risk (stethoscopes) need third-party lab testing; high-risk diagnostic products affecting clinical decisions require real-world testing and sometimes clinical trials.

- The editorial argues generative AI medical products should generally face more comprehensive scrutiny.

Proposed standard:

- Pre-registered clinical trials should become standard for AI in health/medicine (as with drugs and vaccines).

- Not every device needs full-scale RCTs, but the principles must apply: data transparency, open data, and real-world testing are non-negotiable.

The self-driving car parallel: Governments already take a cautious, drug-like regulatory approach to driverless cars — thousands of hours of lab, simulated, and supervised street testing — because errors are catastrophic. The editorial argues AI medical devices deserve the same rigor.

The FDA’s August 2026 discussion paper proposes a two-axis risk framework for generative AI medical devices—assessing both how directive the tool’s outputs are and the potential severity of harm from errors. The agency is considering a “competency-based” premarket evaluation combining non-clinical benchmarking with clinical confirmation, inspired by how physicians are trained and credentialed, rather than relying solely on traditional software testing. A notable proposal is the Foundation Model Device Master File, which would allow foundation model developers to submit confidential model-level information directly to FDA, letting device sponsors reference it without public disclosure.

#aiinmedicine #FDA #clinicaltrials #medicaldevices #Regulation


Tools that inform clinical decision-making require rigorous assessments equivalent to processes used to approve drugs and self-driving cars.

Leave a Comment

Lifeboat Foundation respects your privacy! Your email address will not be published.

/* */