Contact

Health

FDA's generative-AI device paper: a risk grid, not yet a rulebook

The FDA's non-binding paper would weigh risk by how independently a tool acts and how badly it fails, and asks who owns model updates.

Beams of light form a grid across a dark floor, warming to gold in the far corner, as small glowing AI agents hover above and a small figure watches from the near corner.

On 18 August 2026 the US Food and Drug Administration set out, at length, how it might approach medical devices built with generative AI. The document, Considerations for the Regulation of Generative AI-Enabled Medical Devices, comes from the Digital Health Center of Excellence inside the agency's Center for Devices and Radiological Health. FDA is asking for feedback by 19 October 2026 under docket FDA-2026-N-7874.

Its status needs stating plainly. This is a discussion paper, not draft or final guidance. The agency says it neither proposes policy changes nor communicates what evidence it will expect in future submissions, and it deliberately avoids the question of whether FDA's current legal powers cover the ideas it describes. Nothing in it binds a manufacturer, and the approach may change once comments are in.

It still matters. A Department of Health and Human Services spokesperson told MedTech Dive, which reported on 9 September, that FDA has authorised more than 1,500 devices with an AI component but not yet a generative AI-enabled device. The paper shows which questions the agency is weighing for that first applicant: how risk is graded, how an open-ended system is tested, and who answers for a foundation model that someone else updates. Those questions apply to any team putting a conversational assistant or agent near patients, whether or not it ever meets a US regulator.

Why generative systems strain device review

Conventional software review leans on testing a representative sample of inputs against expected outputs. That works when inputs are bounded and outputs are fixed. The paper argues that generative devices break both assumptions: they accept open-ended input, can carry out several subtasks, and may answer similar questions differently. They can also change after release through the model itself, the prompts, the retrieval strategy, guardrails, orchestration logic or the interface.

Many such products sit on general-purpose foundation models from third parties, which disclose varying amounts about training data, architecture and evaluation. That makes it hard to say whether an error belongs to the device or to the model beneath it. The agency's list of resulting risks includes fabricated outputs that look credible, blurred edges around intended use, poor visibility into those underlying models, and performance that degrades across settings and over time.

Two framing points matter. First, FDA regulates devices rather than generative AI as a technology, and its oversight attaches to individual software functions; many software functions are not devices at all. Second, the paper builds on a Digital Health Advisory Committee meeting held on 20 and 21 November 2024, where the agency had already flagged that new evaluation methods might be needed and that premarket evidence might need to be complemented by monitoring after launch.

A grid for risk: what the tool does and what a wrong answer costs

The centrepiece is a two-axis framework, which FDA offers as a possible way to organise thinking rather than as a rule. One axis captures what a function does and how independently it does it. The other captures how much harm could follow if someone relies on an incorrect output. Risk rises from the corner where a tool offers low-stakes, non-directive information towards the corner where it acts on its own in high-stakes situations.

Along the activity axis, the law firm Maynard Nexsen reads four broad positions: non-directive information, information that directs an action, action under supervision, and fully autonomous action. The paper treats the line between informing and directing as a continuum. Its illustration moves from general information, to relating common practice to the user's situation, to endorsing a step, to giving a precise instruction. FDA says it is weighing whether directiveness depends on the substance and context of an output rather than on its wording alone, and whether a line advising a patient to consult their doctor makes an instruction any less directive. Adding that kind of caution, the paper suggests, leaves the underlying instruction just as directive as it was before.

In practice, a disclaimer does not change what a reply tells someone to do.

Functions that act for themselves, such as writing a prescription or starting a set of clinical orders, are presumed riskier than informational ones, although a low position on the consequence axis can temper that. The audience matters too. Patient-facing functions may move up the consequence axis because patients are less able to spot a wrong answer, though the paper also credits the value of patient access to clinical information. Outputs that users cannot check for themselves, such as measurement or signal-processing results, may move up as well.

Two further ideas matter for conversational products. Risk would be judged across realistic conversations, because an exchange that begins with general information can drift into direction. And for functions that advise whether to seek emergency care, both failing to escalate and escalating needlessly count as errors. The first discussion question asks whether other dimensions, such as reversibility, downstream safeguards, time pressure or traceability to source material, belong in the model.

Testing a device the way clinicians are assessed

Because no one can test every possible input, FDA is considering a competency-based approach loosely modelled on how clinicians are examined, supervised and reassessed. It has two parts, non-clinical benchmarking and clinical confirmation, and it would evaluate the final user-facing device as deployed, not the foundation model on its own.

Benchmarking is described as ten possible elements in four groups. Safety covers recognising safety-critical situations and escalating, staying within scope, and conveying uncertainty while deferring to a clinician when needed. Clinical proficiency covers knowledge and task fidelity, gathering and analysing information, quantitative reasoning and communication. Generalisability covers reproducibility across repeated runs and rephrasings, and consistent performance across subgroups, dialects and literacy levels. A final element applies only to agentic devices, and not every element would apply to every product.

The ground rules are familiar: fix methods and acceptance criteria before testing, ground scoring rubrics in clinical guidelines or qualified expert input, and use adjudicators who are independent of both the sponsor and the model developer. FDA adds that the independence test would still apply if the adjudicator were itself a language model. It also concedes that public benchmarks can be contaminated, saturated or unrepresentative of real use.

Clinical confirmation would not always require a prospective study. The paper lists options in rough order of rigour and patient exposure: retrospective runs on real patient inputs, shadow deployment in a live workflow where outputs reach no one, scripted sessions with trained patient actors, independent clinician review of real cases, and prospective studies, in some cases randomised trials. For open-ended outputs with no single right answer, performance might be compared with a panel of qualified clinicians or with a median clinician in practice. FDA also asks whether independent third parties might hold sequestered test sets or run parts of the assessment.

After launch: monitoring, and models that change underneath

The weightiest idea may be a trade between premarket and postmarket evidence. FDA is asking whether it could accept more uncertainty at authorisation if a device comes with strong monitoring once in use, and for which kinds of device that would be inappropriate. Maynard Nexsen describes this as the most consequential question in the paper.

The monitoring options sketched are periodic re-benchmarking on a set schedule and after trigger events such as an update to the underlying model, sampled review of real interactions by independent clinicians, and tracking of drift against defined thresholds. The agency also asks whether automated supervisory agents could carry part of that load, and how the reliability of such agents would itself be judged.

Change control is where third-party models bite. The paper separates deliberate changes by the sponsor, gradual changes in systems designed to adapt, and unplanned changes that arrive when a foundation model developer updates its model. Oversight could range from a record in the quality management system to FDA approval before release. A predetermined change control plan, a mechanism FDA already uses for AI-enabled device software, is offered as one route, with the premarket benchmark serving as the baseline for re-testing after each modification. Question 24 asks directly how manufacturers can detect and respond to changes they did not start, whether through contracts, technical means or such a plan.

A master file for foundation models, and the agentic question

To make review of devices built on the same model more consistent, FDA floats a voluntary Foundation Model Device Master File, drawing on its existing master file programme. Model developers and platform providers could lodge material such as model or system cards, covering architecture, training data provenance, known limitations and failure modes in healthcare settings, subgroup results, built-in guardrails, commitments to give notice of updates, and the availability of audit logs. FDA would hold the file confidentially, and a sponsor could cite it with the holder's permission.

The limits are explicit. Filing would not authorise the model for any device use, and each sponsor would still have to show that its own device is safe and effective. The agency raises the obvious difficulty itself in question 25: model developers may have little incentive to disclose safety-relevant information voluntarily, or to keep a file current.

Agentic systems receive shorter treatment. FDA defines them as generative systems that plan and carry out multi-step tasks, use external tools or act across a sequence of steps. It notes they are appearing in care coordination, documentation, patient outreach and workflow support, much of which may fall outside device oversight, but says an agent whose actions end up controlling another medical device could meet the definition of a device. The proposed agentic benchmark probes planning within intended limits, recognition of faulty tool outputs, respect for human checkpoints before irreversible or high-consequence steps, and resistance to prompt injection arriving through user input, retrieved content or tool outputs.

What is disputed or still unknown

The legal footing is the first open question. The paper does not say whether FDA's existing authorities stretch to what it describes, and the eventual framework, including the master file, may look quite different after comments. Arnold & Porter notes that the paper leaves FDA's current clinical decision support policy unchanged and that the two frameworks do not line up exactly, which leaves room for arguments at the boundary.

Industry behaviour shows the cost of that uncertainty. Suzanne Levy Friedman, a partner at Honigman, told MedTech Dive that most generative tools now on the market steer clear of regulated device functions and favour administrative uses until FDA spells out its requirements. 'No one really wants to be the guinea pig,' she said.

Evidence quality is another fault line, and FDA says so. It asks how a sponsor can show that benchmark scores predict real-world behaviour, how much weight sponsor-built benchmarks deserve given the risk of tuning to the test, and whether synthetic data produced by models of the same class could hide the very gaps an evaluation is meant to expose.

Responsibility for monitoring is also unresolved. Kellie Owens, an assistant professor of medical ethics at NYU Grossman School of Medicine, told MedTech Dive that vendors should keep their products performing well but that in her experience this does not always happen. Kayla Cristales of Haynes Boone expects FDA may end up relying heavily on companies to supply updates. The paper's own question 21 asks how clinicians, healthcare institutions and standards bodies could share the monitoring burden without diluting the manufacturer's accountability.

Practical implications for teams deploying AI

None of this is binding. It still describes a sound discipline for any organisation running AI, voice agents or automation in health-adjacent operations, in the US or elsewhere.

  • Place every function on both axes. Record how directive or autonomous each output is and what a wrong output would cost, test whole conversations rather than single replies, and do not count disclaimers as mitigation.
  • Benchmark the system you actually ship. Treat prompts, retrieval, guardrails and orchestration as one versioned configuration, set acceptance criteria before testing, measure failures in both directions, such as missed and needless escalations, and use reviewers who did not build it.
  • Treat an upstream model update as a change event. Pin model versions where the supplier allows it, negotiate advance notice of updates in contracts, and re-run benchmarks before a new version reaches users.
  • Put checkpoints in front of irreversible actions. Agents should pause for a human before high-consequence steps, and prompt-injection testing should cover retrieved documents and tool outputs, not only user messages.
  • Budget for monitoring from the start: sampled human review of real interactions, drift thresholds with named owners, and a written split of duties between the vendor and the organisation deploying the tool.
  • If the discussion questions touch your product, particularly those on third-party model changes, master files and agentic systems, consider commenting before 19 October 2026. Specific, evidence-backed submissions give the agency more to work with than general endorsements.

Sources

  1. Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for FeedbackU.S. Food and Drug Administration · 18 August 2026
  2. FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical DevicesU.S. Food and Drug Administration · 18 August 2026
  3. FDA Releases Discussion Paper on Generative AI-Enabled Medical Devices: What Manufacturers and Investors Need to KnowMaynard Nexsen · 26 August 2026
  4. FDA Seeks Public Feedback on Regulatory Approach for Generative AI-Enabled Medical DevicesArnold & Porter · 1 September 2026
  5. 4 questions about the FDA's approach to generative AIMedTech Dive · 9 September 2026