Generative AI Medical Device Risk Assessment
Generative AI medical device risk assessment should evaluate what the device can influence, what happens when its output is wrong, how variable its behavior can be, and how much human review exists before the output reaches care. The product label alone is not enough.
FDA's August 2026 discussion paper raises a possible two-axis approach to risk assessment, but the agency is explicit that the paper is for discussion and does not establish policy. Health systems can still use the questions behind it now: clinical consequence and system behavior both matter.
Start With Intended Use
Risk assessment begins with the intended use, users, patient population, inputs, outputs, and care setting. A generative system that drafts low-risk administrative text creates a different exposure than a system that generates diagnostic content, recommends treatment, or takes actions through connected tools.
Do not classify the product only by the model type. The same foundation model can support very different functions. The clinical claim and deployed workflow determine the practical risk.
Clinical Consequence
Ask what could happen if the output is wrong, incomplete, delayed, or persuasive but unsupported.
Consequence may include:
- missed diagnosis
- unnecessary testing
- treatment delay
- privacy exposure
- incorrect documentation
- a failure to escalate urgent information
Severity is only one part of the review. Teams also need to consider how often the function is used, whether errors are easy to detect, and whether one failure can affect many patients or downstream systems.
Output Variability
Generative AI can produce different outputs from similar inputs.
That variability may be useful, but it complicates repeatable testing.
Risk assessment should cover:
- realistic prompts
- incomplete data
- conflicting context
- uncommon cases
- attempts to push the system beyond its intended use
The test set should include the messy inputs the device will actually receive. Clean benchmark prompts are not a substitute for clinical notes with ambiguity, copied text, missing information, and local abbreviations.
Autonomy and Connected Actions
Agentic functions increase risk when a system can choose tools, sequence steps, retrieve records, place information into another system, or initiate an action with limited review. The assessment should identify every action the system can take and every control that limits it.
A tool that generates a suggestion for a clinician to review is different from a tool that sends a message, changes a queue, or writes into the medical record. Human review has to occur before the point where an error becomes difficult to reverse.
Foundation Model Dependencies
A medical device may depend on a foundation model supplied by another company. That creates version, availability, data-use, and change-control questions. The manufacturer and health system should know which model is used, how updates are managed, what behavior can change, and which changes trigger revalidation.
Dependencies also include:
- retrieval systems
- terminology services
- EHR data
- calculators
- connected applications
A model can appear stable while one of those inputs changes underneath it.
Human Review Is a Control Only When It Is Real
Calling a system human in the loop does not settle the risk question. The review step must be specific. Who reviews the output? What information do they see? How much time do they have? Can they identify unsupported content? What happens when they disagree?
Review can fail when users are overloaded, when the generated answer looks polished, or when the workflow makes correction harder than acceptance. Training and interface design are part of the control.
A Practical Risk Record
- intended use, users, and clinical setting
- patient population and data inputs
- generated outputs and connected actions
- credible failure modes and clinical consequences
- required human review and escalation rules
- foundation model and external dependencies
- validation evidence and known limits
- monitoring measures and stop criteria
- version-change and re-review triggers
Risk Assessment Is Not Approval
The assessment does not decide by itself whether the device should be deployed.
It gives the governance group enough information to choose the right next step:
- reject
- request evidence
- test locally
- pilot with controls
- approve for limited use
- require more monitoring
The useful output is a decision record that can be revisited when the model, workflow, or evidence changes.
Related AI Medicine Now Coverage
- Regulated Generative AI Medical Devices
- Postmarket Monitoring for Generative AI Medical Devices
- Regulatory Readiness
- Clinical AI Governance Framework
- AI Product Release Criteria for Clinical Tools
Reviewed: September 2, 2026. Next review: October 20, 2026.
Frequently Asked Questions
What should a generative AI medical device risk assessment cover?
It should cover intended use, clinical consequence, output variability, autonomy, connected actions, human review, model dependencies, validation, monitoring, and change control.
Has FDA finalized a risk framework for generative AI medical devices?
No. FDA's August 2026 discussion paper seeks public input and does not establish draft guidance, final guidance, or policy.
Does human review automatically make a generative AI medical device low risk?
No. Review is useful only when the reviewer has the time, information, training, authority, and workflow support needed to detect and correct errors.
Related Reading
Regulated Generative AI Medical Devices: What the FDA Discussion Paper Means
FDA's August 2026 discussion paper does not create new policy, but it shows how the agency is thinking about risk, competency assessment, monitoring, foundation models, and agentic AI in medical devices.
Postmarket Monitoring for Generative AI Medical Devices
Postmarket monitoring for generative AI medical devices should track output errors, corrections, clinical escalation, subgroup performance, dependencies, and model changes.
Clinical AI Governance Framework
A clinical AI governance framework gives hospitals a way to review, deploy, monitor, and retire AI tools with clear accountability. The goal is not bureaucracy for its own sake, but safer decisions around risk, evidence, privacy, workflow, vendor management, and ongoing oversight.
AI Product Release Criteria for Clinical Tools
AI product release criteria help health systems and vendors decide whether a clinical AI tool is ready for pilot, go-live, expansion, or re-release after a model or workflow change.
Clinical AI Inventory: How Health Systems Track AI Tools
A clinical AI inventory helps health systems know which AI tools are in use, who owns them, what data they touch, what evidence supports them, and what monitoring is required after deployment.