On August 18, 2026, the U.S. Food and Drug Administration’s (“FDA”) Digital Health Center of Excellence, within the Center for Devices and Radiological Health (“CDRH”), released a discussion paper seeking public comment on how the agency should approach generative artificial intelligence-enabled (“GenAI”) medical devices. FDA emphasized that the paper is intended for discussion purposes only: it does not represent draft or final guidance, propose or implement policy changes, communicate proposed or final regulatory expectations or address whether the approaches discussed would fit within FDA’s existing statutory authority.
The paper is relevant only where a GenAI-enabled function meets FDA’s definition of a medical device; FDA does not regulate a technology as a medical device merely because it includes AI. Even within that boundary, the paper is significant because it highlights practical questions FDA is examining as GenAI-enabled devices become more capable, conversational and harder to evaluate through traditional software validation approaches, including questions about risk, evidence, monitoring and change control.
The paper is also noteworthy because FDA’s prior AI regulatory efforts have largely focused on predictive or adaptive machine learning systems. This discussion paper is one of FDA’s most comprehensive discussions to date regarding how generative AI capabilities, including conversational interfaces, foundation models and agentic workflows, may challenge existing medical device regulatory frameworks.
Why GenAI-Enabled Devices Raise Different Regulatory Questions
FDA identifies several characteristics that may distinguish GenAI-enabled devices from traditional software and earlier AI-enabled devices. Unlike software with bounded inputs and predictable outputs, GenAI-enabled tools may accept open-ended inputs, perform multiple subtasks and generate different outputs in response to similar prompts. They may also change over time through updates to the underlying model, prompts, retrieval strategies, guardrails, orchestration logic or user interface.
FDA recognizes that these capabilities may offer benefits, including more personalized outputs, improved adaptability to novel inputs, enhanced user interaction and support for more complex decision-making. At the same time, they may complicate risk assessment and evidence development. FDA identifies potential risks, including hallucinations, unclear intended-use boundaries, limited transparency into underlying models and performance degradation across testing and real-world deployment. For products built on third-party foundation models, sponsors may face additional challenges if they lack visibility into model training, architecture, evaluation methods, limitations or updates. Foundation models are general-purpose AI systems that can be adapted to perform a broad range of tasks, including health care-related functions for which they were not originally developed.
FDA Explores a Risk Assessment Framework
FDA’s risk framework discussion starts with two questions: What is the GenAI-enabled device function doing, and how serious could the consequences be if the output is wrong? FDA seeks input on a possible two-axis framework for assessing risk in GenAI-enabled device functions. One axis would focus on whether the function provides information, directs action or takes action with varying degrees of human oversight. The other would focus on the severity of harm that could result if a user relies on an incorrect output.
In general, risk would increase as a function moves from providing general information to directing or taking action, and as the potential consequences of error become more serious. FDA emphasizes that “directiveness” may exist on a continuum. An output may become more directive as it becomes more specific, personalized or action-oriented, even if it avoids words such as “recommend,” “should” or “consider.” FDA also notes that a disclaimer such as “talk to your doctor” may not make an otherwise action-directing output less directive.
FDA also asks how the framework should apply in more complex settings, including patient-facing tools, tools used by generalists rather than specialists, multi-turn conversations that shift toward recommendations or instructions and care-escalation functions. FDA also highlights measurement, signal-processing and in vitro diagnostic functions, in which users may not be able to independently assess whether an output is correct. For care-escalation tools, FDA is considering both under-escalation and over-escalation. Under-escalation may delay needed care, while over-escalation may lead to unnecessary care, resource burden, patient anxiety and reduced trust over time.
Premarket Evaluation
Because GenAI-enabled devices may produce open-ended outputs across a wide range of scenarios, FDA asks whether traditional software testing should be supplemented by a competency-based evaluation model. FDA analogizes the concept to competency-based assessments used in health care professional training and credentialing, with structured evaluation of core competencies followed by confirmation in real or representative practice settings.
For GenAI-enabled devices, FDA describes an approach that would focus on the final device as it would be deployed, rather than the foundation model alone. The approach would pair device benchmarking with clinical confirmation, with the type and amount of evidence tailored to the device’s intended use and risk profile.
Benchmarking could assess whether the device demonstrates relevant safety behaviors, clinical proficiency, communication quality, generalizability, and, for agentic systems, the ability to safely plan and execute multistep tasks. FDA contemplates that sponsors would prespecify the scope of testing, evaluation methods, scoring rubrics, expert adjudicators and acceptance criteria.
Because benchmarking may not fully predict performance in clinical use, FDA also discusses clinical confirmation under real or clinically representative conditions. Depending on the device and its risk profile, possible approaches could include retrospective evaluation using real patient inputs, shadow deployment, standardized patient interactions, clinician adjudication of real cases or a prospective clinical study. FDA does not suggest that every GenAI-enabled device would require a prospective trial.
Foundation Model Device Master Files
One of the most notable concepts in the discussion paper is FDA’s exploration of Foundation Model Device Master Files (“MAFs”). If ultimately adopted, the approach could create a mechanism through which foundation model developers participate more directly in the medical device regulatory ecosystem without becoming the sponsor of a specific device.
The paper also addresses devices built on third-party foundation models. These models may influence refusal behavior, content policies, output formatting, version control and other safety-critical characteristics, while device sponsors may have limited access to proprietary information about the model.
To address that gap, FDA seeks comment on whether foundation model developers and platform providers could use FDA’s existing Device Master File program to submit voluntary Foundation Model MAFs. Those files could contain confidential model-level information submitted directly to FDA. With the file holder’s authorization, device sponsors could reference that information in their own premarket submissions, allowing FDA reviewers to consider model-level information without requiring public disclosure or direct disclosure to each sponsor.
FDA suggests that a Foundation Model MAF could include information relevant to FDA’s review, such as model architecture, training data provenance, supported use cases, known limitations, failure modes, health care-relevant benchmark results, subgroup performance, safety guardrails, update commitments and audit-log availability. This could help reduce duplicative submissions and support more consistent review when multiple device sponsors rely on the same foundation model.
Importantly, a Foundation Model MAF would not authorize or clear the foundation model for any medical device use. Device sponsors would remain responsible for demonstrating the safety and effectiveness of their own device, including its configuration, intended users, intended use and deployment context.
Post-Market Monitoring and Change Control
Because GenAI-enabled devices may produce variable outputs and may change after deployment, FDA asks whether stronger post-market monitoring could, in some circumstances, support greater tolerance for premarket uncertainty. The paper identifies several possible monitoring approaches, including periodic re-benchmarking, sample-based review by independent clinicians and monitoring for performance degradation or drift. FDA also asks whether machine-based supervisory agents could support monitoring, while recognizing that those supervisory tools would also need to be evaluated.
FDA also discusses how post-market monitoring could support change control. A device’s premarket competency assessment could serve as a baseline for evaluating later modifications, including updates to the model, prompts, retrieval strategy, guardrails or other components. After a change, the sponsor could re-benchmark the device against the same capabilities to assess whether safety or effectiveness may be affected. FDA also notes that Predetermined Change Control Plans (“PCCPs”), which allow sponsors to describe certain planned device modifications in advance, may help manage certain changes, although GenAI-related modifications may be difficult to fully prespecify.
The discussion highlights a potential tension between PCCPs and GenAI systems. PCCPs are designed to permit certain predefined modifications without additional submissions, but GenAI-enabled systems may evolve in ways that are difficult to fully specify in advance. This raises questions about how PCCPs could accommodate foundation-model updates, retrieval changes, prompt-level modifications or other changes that may affect device performance.
Agentic AI Systems
FDA separately asks how to evaluate agentic AI systems, which are GenAI-enabled systems that can plan and execute multistep tasks, use external tools or take actions across a sequence of steps. Some health care uses, such as care coordination, clinical documentation, patient outreach or workflow support, may fall outside FDA’s device jurisdiction. However, when an agentic function meets the device definition, such as where it controls another medical device or takes clinically significant action, FDA suggests that additional evaluation questions may arise. These questions include how autonomy, tool use, human oversight and chains of related actions should affect acceptance criteria and oversight.
Practical Considerations for Developers, Sponsors and Organizations Deploying or Using GenAI Tools
Although FDA is not proposing binding requirements, the discussion paper identifies issues that may shape future expectations for GenAI-enabled medical devices. Developers, sponsors, foundation model providers and organizations deploying or using these tools should consider how FDA’s focus on risk, evidence, monitoring, change control and model transparency may affect product development, regulatory strategy, vendor oversight, clinical deployment and governance.
- Assess whether the GenAI-enabled function meets the medical device definition. Companies should consider not only intended-use statements and product claims, but also realistic outputs, user interactions, clinical context and whether the tool directs or takes action.
- Pressure test risk classifications. For products that may fall within FDA’s device jurisdiction, companies should evaluate risk in light of directiveness, autonomy, patient-facing use, generalist versus specialist users, care escalation and the consequences of an incorrect output.
- Revisit validation and evidence plans. Sponsors should assess whether their existing validation approach is sufficient for GenAI-enabled functionality, particularly where the tool may produce variable outputs, operate across a broad range of scenarios or require confirmation in real or clinically representative settings.
- Review dependencies on third-party foundation models. Companies should assess contractual and technical controls for transparency, versioning, audit rights, update notifications, service commitments, incident response and contingency planning. A future Foundation Model MAF could support FDA review, but it would not replace sponsor responsibility for the finished device.
- Plan for deployment, monitoring and provider governance. Companies and organizations using GenAI-enabled tools in clinical workflows should consider governance processes for vendor selection, local validation, output review, user training, escalation pathways, clinical oversight, incident reporting, performance review and coordination with manufacturers, particularly if FDA places greater emphasis on post-market evidence and monitoring.
- Consider whether to comment on FDA’s specific questions. Because FDA is seeking input across multiple areas, stakeholders may want to submit comments on issues that could affect their products or operations, including risk classification, evidence expectations, post-market monitoring, PCCPs, foundation model transparency and evaluation of agentic AI systems. Comments are due by October 19, 2026, and may be submitted through Regulations.gov under Docket No. FDA-2026-N-7874.
Taken together, the discussion paper signals that FDA is looking beyond the evidence needed before marketing. The agency is also focused on how GenAI-enabled devices are governed, monitored and updated over time. Stakeholders developing, deploying or relying on these technologies should consider whether to engage in the comment process, particularly where FDA’s questions could affect product design, evidence strategy, foundation model relationships, post-market monitoring or clinical deployment.
For further information or assistance regarding this topic or assistance submitting comments, please contact:
- Carolina Wirth at (202) 780-2989 or cwirth@hallrender.com;
- Melissa Markey at (248) 740-7505 or mmarkey@hallrender.com; or
- Your primary Hall Render contact.
Hall Render blog posts and articles are intended for informational purposes only. For ethical reasons, Hall Render attorneys cannot give legal advice outside of an attorney-client relationship.