Summary 🔗
On August 18, 2026, CDRH's Digital Health Center of Excellence released Considerations for the Regulation of Generative AI-Enabled Medical Devices, a discussion paper and request for feedback under docket FDA-2026-N-7874. Comments are due October 19, 2026. The paper is not draft or final guidance, proposes no policy changes, and does not address whether its approaches fit within FDA's existing legal authority. It contains 26 numbered discussion questions.
The paper proposes a two-axis risk framework. One axis scores how independently a function acts, from non-directive information through action-directing information to supervised and fully autonomous action. The other scores the consequence of relying on an incorrect output, from limited to severe. Evidence expectations scale with position on the grid, not with the presence of an LLM.
Premarket evaluation follows a competency-based approach modeled on how clinicians are credentialed. First, device benchmarking: high-throughput non-clinical testing of the deployed configuration across ten elements spanning safety, clinical proficiency, generalizability, and agentic conduct. Second, clinical confirmation through five approaches of increasing rigor: retrospective evaluation, shadow deployment, standardized patient interactions, clinician adjudication, and prospective studies. Not every device would require a prospective study.
Postmarket, FDA asks whether it should accept greater premarket uncertainty in exchange for continuous monitoring: periodic re-benchmarking, sample-based clinician review of real-world outputs, and drift detection. Section VII floats a voluntary Foundation Model Device Master File through which model developers could confidentially share architecture, training data provenance, and update commitments that sponsors reference with permission.
Postmarket monitoring: the trade I asked for at DHAC 🔗
Most of FDA’s paper is par for the course for anyone who builds regulated AI: a risk grid, benchmarks and evals before clinical deployment, a ladder of clinical confirmation, all covered in the sections below. Solid work, and a signal that FDA will entertain pragmatic regulation. One sentence on page 19 is what I think to be the highest impact thought though: "CDRH is considering whether it is appropriate to accept greater premarket uncertainty regarding a GenAI-enabled device's benefit-risk profile through greater reliance on postmarket monitoring."
Premarket studies and postmarket monitoring are both risk controls, both there to verify that the device is safe, effective, and will generalize to the U.S. population. The real question in Section VI is which control does that job better.
I think the answer is lopsided. A premarket study measures a snapshot: one model version, one test population, frozen months before launch, funded before a dollar of revenue. Premarket evidence can bankrupt companies before they get to market. Postmarket monitoring catches failures that are impossible to predict with premarket evidence alone. Any serious manufacturer will run it after v1.0 ships anyway. Nobody has a crystal ball premarket; a monitoring pipeline does not need one. Shifting evidence burden from before market to after it is the rare policy change that saves sponsors money and time while handing FDA more evidence, not less. I do not know of another lever in this framework that does both.
The monitoring options are periodic re-benchmarking against the premarket thresholds, sample-based review of real-world inputs and outputs by independent clinicians, and drift monitoring. FDA asks whether "machine-based supervisory agents" can run some of it (Question 20). The premarket assessment becomes the baseline: a modified device is "re-benchmarked against the same capabilities," with a PCCP as one mechanism, and third-party foundation model updates get their own question (Q24) because the vendor rather than the sponsor may trigger the change.
On November 21, 2024, at FDA's first Digital Health Advisory Committee meeting on generative AI, I used my five minutes in the open public hearing to propose that trade: a frontier model generating site-specific synthetic test data, automated ground-truthing, and a nightly comparison against device outputs so drift is caught as it happens, offered as the reason FDA could accept a lighter premarket package. In the Q&A I called the PCCP "test-driven development, but you're getting your tests pre-approved." The pre-submission I had published a month earlier asked FDA whether nightly reruns were enough to detect a cloud vendor changing the model without notice. I restated the 90/10 premarket-to-postmarket split in February. FDA's paper adopts none of it as policy, but it asks the industry whether it should, and a question on the docket is how policy changes begin. My position has not moved: trading premarket burden for postmarket rigor is better for industry and FDA alike. It converts capex to COGS, and the safety checks keep running long after the review checkpoint.
I have been advocating for LLM-as-a-judge for years 🔗
The LLM adjudicator sentence is worth dwelling on because it can significantly reduce the burden of proof by leveraging expensive human reviewers more effectively. So rather than manually annotate 10k reports, annotate 1k to train an LLM to judge from that point onwards. Cliche saying (sorry) but its like teaching a computer how to fish instead of giving it a fish.
I have been asking FDA to accept that judge since 2024. My October 2024 pre-submission asked, in writing, whether an LLM that is not the device model could adjudicate a secondary test set, and it included a methodology I called Automated ML Verification Verification: a judge for the judge, with engineers and radiologists auditing the LLM grader against acceptance rates above 95 percent and a second frontier model checking the first. This was before the concept of LLM-as-a-judge entered the common vernacular. Question 8 of my pre-sub asked whether the methodology was sufficient for proactive postmarket surveillance. FDA's paper now puts the same questions to the whole industry.
I have been arguing for this framework since 2024 🔗
FDA built this paper on the record of that November 2024 DHAC meeting and cites it as its origin. The postmarket trade was not the only position I put on the record at the microphone.
That DHAC exchange was not a one-off. In July 2024 I published a foundation model FAQ that flagged benchmark contamination before FDA gave it a section number. In October I published a pre-submission that asked FDA, in writing, whether an independent LLM could adjudicate a test set and whether nightly re-benchmarking would catch a cloud vendor changing the model unannounced. In November I argued the postmarket trade in person at DHAC, and in December, from the RSNA podium, I told sponsors to constrain the shell so a user cannot turn a medical device into ChatGPT. FDA might have been reading along the whole time.
The July receipt is the most literal. My foundation model FAQ told sponsors to check their test data prospectively, because a foundation model's sprawling training set can silently swallow a public benchmark. FDA's paper now names the same failure modes for publicly available benchmarking assets, in the same words: data contamination, saturation, and lack of representativeness.
In an October 2024 pre-submission, I surfaced the nightly-rerun idea into the record. Question 6 of that document proposed a standalone performance test repeated nightly to catch a third-party cloud ML provider changing the model unexpectedly, and asked FDA whether that was sufficient to show stable configuration management. FDA’s 2026 paper's change-control discussion now flags exactly that scenario, changes initiated by the foundation model developer rather than the device manufacturer, and hands it to the whole industry as Question 24.
December closed the year at RSNA, where I told sponsors to constrain inputs and outputs so a user cannot turn a cleared device into ChatGPT, and to test against prompt injection and jailbreaking attacks. That advice now has a benchmark element number: S.2, scope maintenance and boundary adherence, which probes adversarial prompting, prompt injection, and multi-turn drift out of scope, and counts over-refusal as a failure too.
What FDA published, and what it is not 🔗
The paper is 30 pages with 26 numbered discussion questions, written by DHCoE under Director Rick Abramson and CDRH Director Michelle Tarver, with Acting Commissioner Kyle Diamantas framing it as part of the push to move AI medical products to market faster. It grows out of the November 2024 DHAC meeting on generative AI that I participated in person, which the paper cites as its origin.
Read the disclaimer before the hype. The paper "does not represent draft or final guidance," is "not intended to propose or implement policy changes," and does not address whether the approaches fit within FDA's existing legal authority. RAPS's same-day roundup called it draft guidance. It is not. It is an early look at the evidence structure FDA is considering, offered before the agency commits, so the comment period is where industry negotiates that structure. FDA also restates the principle that governs the rest: it "does not regulate GenAI as such; it regulates medical devices."
What is FDA's two-axis risk framework for generative AI devices? 🔗
FDA's two-axis framework scores a generative AI function by what it does and by the consequence of relying on an incorrect output. The activity axis runs from non-directive information (a cardiovascular risk score) through action-directing information (a pointed push to seek emergency care) to action-taking under clinician supervision, and finally to fully autonomous action. The consequences axis runs from limited to severe. Risk rises from lower left to upper right.
FDA's examples show where its attention is. Hydrocortisone for poison ivy sits low. Whether to go to the ED for chest pain, or how to adjust basal insulin, sits high even though both are "just information." Autonomously prescribing antibiotics for confirmed strep is lower than autonomously starting a thrombolytic order set for stroke. The evidence expectation follows the position on the grid rather than the presence of an LLM.
Section IV's finer points will reshape intended-use statements. Directiveness is a continuum that FDA judges on substance, and adding "talk to your doctor" does not make an action-directing output less directive. Patient-facing functions may move up the consequences axis because patients cannot check the output. FDA assesses multi-turn conversations on realistic trajectories, since a chat that starts informational can drift into directing action. And reviewers score escalation in both directions: sending everyone to the ED counts as harm.
Your intended-use sentence now has to say how independently the function acts and how bad a wrong answer is, and you have to defend both placements in your risk file. That is the line I drew in my UpDoc analysis in June. FDA had cleared conversation on the outside and a deterministic protocol on the inside, not an autonomous LLM physician.
What is FDA's competency-based approach for generative AI devices? 🔗
FDA's competency-based approach evaluates a generative AI device the way medicine evaluates a physician: standardized benchmarking of knowledge, safety behavior, and generalizability, then clinical confirmation in settings of increasing realism, then ongoing assessment in use. Clinicians are not tested on every scenario they might meet, FDA reasons, and neither can a device with open-ended inputs be. The paper cites Patel and Blumenthal (JAMA Health Forum, 2026), Bergman, Wachter, and Emanuel's licensure framework (JAMA, 2026), and Freyer et al. (Nature Medicine, 2025), then adapts the idea to device law.
The structure is familiar to anyone who tracks frontier models. Artificial Analysis and the other public leaderboards grade each new LLM across separate competencies: math, science, coding, spatial reasoning, even humor. FDA is proposing the same decomposition for medical devices, with the skill list swapped for safety behavior, clinical proficiency, generalizability, and agentic conduct. Centralization is the detail I keep coming back to: FDA is entertaining shared benchmarks in the style of Artificial Analysis, a common exam rather than a bespoke test set per sponsor, and a common exam gives reviewers scores they can compare across submissions.
Build your architecture around two sentences in Section V. The evaluation target is "the final user-facing device, as configured and intended to be deployed for real-world use, and not the foundation model standing alone." And the approach is "proportionate to its risk," with the two-axis grid setting how much benchmarking and confirmation.
Device benchmarking: the ten elements 🔗
Benchmarking is FDA's word for high-throughput, non-clinical testing of the deployed configuration. The paper proposes ten elements, chosen per device by intended use and risk:
| Group | Element | What FDA wants probed (Appendix A) |
|---|---|---|
| Safety | S.1 Safety-critical recognition and escalation | Time to escalation in evolving encounters; resistance to over-reassurance; under- and over-escalation |
| Safety | S.2 Scope maintenance and boundary adherence | Adversarial prompting, prompt injection, emotional manipulation, multi-turn drift out of scope; over-refusal also fails |
| Safety | S.3 Calibration, uncertainty, clinical deferral | False confidence on contested or outdated information is a safety failure |
| Clinical proficiency | E.1 Clinical knowledge and task fidelity | Current guidelines, contraindications, interactions, special populations |
| Clinical proficiency | E.2 Information gathering and analysis | Differential completeness, follow-up questions, premature closure |
| Clinical proficiency | E.3 Quantitative and measurement analysis | Weight- and renal-based dosing, unit conversions, implausible values |
| Clinical proficiency | E.4 Communication and comprehension | Health literacy, empathy, coercive language, automation bias |
| Generalizability | R.1 Robustness, reliability, reproducibility | Repeated runs, paraphrases, input order, long conversations |
| Generalizability | R.2 Subgroup performance | Demographics, dialects, accents, literacy levels |
| Agentic | A.1 Agentic competencies | Planning inside the safety envelope, tool-error recognition, human checkpoints before irreversible actions |
The method principles will be familiar to anyone who has run a reader study: prespecified methods and acceptance criteria, rubrics grounded in guidelines or validated by experts, adjudicators independent of both the sponsor and the model developer. Then comes the sentence I read twice, because it puts in writing that an LLM judge is admissible if it is independent and qualified: those independence expectations "would still be applicable when the expert adjudicator is itself an LLM."
Interestingly, in July I published a 1,200-question regulatory judgment benchmark that graded more than two dozen frontier AI configurations against an opinionated answer key distilled from 15 years of SaMD consulting, with free-text answers scored by a three-judge AI panel on a two-of-three majority and a strict held-out set. That is basically FDA's proposed competency evaluation structure but applied to regulatory consulting instead.
Clinical confirmation: five rungs, not one study 🔗
Benchmarking "may not fully establish how a GenAI-enabled device will perform in real clinical use," so FDA adds clinical confirmation, with explicit relief: it "might not require a prospective clinical study in every case." Five approaches, in rising rigor and patient exposure: retrospective evaluation on real patient inputs (synthetic supplements allowed where data is thin), shadow deployment with outputs recorded but hidden from care, standardized patient interactions with trained actors, blinded or unblinded clinician adjudication of real cases, and a prospective study, sometimes an RCT.
For open-ended outputs with no single correct answer, FDA floats the comparator I have been recommending for generated text: "a panel of qualified clinicians whose consensus reflects the applicable standard of care, or... a median clinician in practice." FDA concedes these designs "may not be powered around traditional effectiveness endpoints" and asks how sample sizes should be set (Question 12). That is an invitation to arrive with a validated LLM-judged error-rate endpoint and a human-adjudicated subsample, sized the way we size sensitivity and specificity studies today.
Foundation models, agentic AI, and the model behind your device 🔗
Section VII floats a voluntary Foundation Model Device Master File: a model developer submits model or system card information that FDA holds in confidence and that sponsors reference with the holder's permission. FDA's footnote lists what it wants: architecture, training data provenance, healthcare failure modes, subgroup benchmark results, guardrails, "update notification commitments," audit log availability. A MAF would not authorize the model for any use; the sponsor still proves the device.
The idea has been in the air for some time. When I ran my first foundation-model pre-submission in 2024, FDA's own minutes suggested the Master File program for a model used as a component of someone else's device, and the day before DHAC I argued that vendors who publish training-data summaries make their customers easier to clear. Question 25 asks the honest follow-up: given "limited incentive to disclose," what would make a voluntary MAF useful?
I think it could be a differentiator that foundation model developers could start publishing in their model cards if they want others to solve the last-mile FDA review problem for them.
What to do before October 19, 2026 🔗
Comment. FDA accepts partial responses, so answer the questions that touch your product: for most LLM device sponsors that is Q7 (is the competency approach right), Q10 (benchmark validity), Q11 and Q12 (confirmation rung and sample size), Q18 (the postmarket trade), Q22 (which changes stay inside the QMS), Q24 (third-party model changes), and Q25 (MAF content). If you only answer one, answer Q18: it decides whether the evidence burden for this category stays parked in front of revenue or moves to where the risk lives. The answers FDA collects here become the evidence bar reviewers grade you against later. Innolitics will file a comment on the docket; if you want a real device scenario represented, anonymized, we will fold it in.
And if you are building one of these devices, this is the work we do. Innolitics is a concept-to-clearance engineering and regulatory firm for healthcare AI and agentic software as a medical device: one team places your functions on the two-axis grid, builds the benchmark and LLM-judge harness, designs the clinical confirmation study, and writes the Pre-Sub, 510(k), or De Novo submission.
Frequently asked questions 🔗
When are comments due on FDA's generative AI discussion paper? October 19, 2026, on Regulations.gov under docket FDA-2026-N-7874. FDA accepts partial responses; you do not need to answer all 26 questions.
Is the FDA generative AI discussion paper binding guidance? No. It is a discussion paper from CDRH's Digital Health Center of Excellence. It is not draft or final guidance and does not propose policy changes or evidence expectations for marketing submissions.
Does FDA regulate LLMs as medical devices? FDA regulates device functions, not models. An LLM-enabled function with a medical intended use is a device. UpDoc's K253281, cleared December 23, 2025, shows a narrow, clinician-supervised LLM function can clear through a 510(k).
What is FDA's competency-based approach? Non-clinical device benchmarking across safety, clinical proficiency, generalizability, and agentic elements, followed by clinical confirmation ranging from retrospective evaluation to a prospective study, with rigor scaled to the device's position on the two-axis risk framework.
What is a Foundation Model Device Master File? A voluntary, confidential filing under FDA's existing Device Master File program in which a foundation model developer supplies model or system card information that device sponsors can reference in their submissions.
Can an agentic AI medical device get cleared in 2026? As a bounded, supervised function with a predicate or a De Novo path, yes. The paper adds an agentic competency element (A.1) and signals heavier scrutiny of multi-step autonomy and tool use.
Can an LLM be used to grade a medical device's test results? Conditionally, yes. FDA's method principles state that adjudicator independence expectations "would still be applicable when the expert adjudicator is itself an LLM." The judge must be independent of the sponsor and the model developer, qualified for the clinical domain, and validated against a human-adjudicated subsample.
Does every generative AI device need a prospective clinical trial? No. The paper says clinical confirmation "might not require a prospective clinical study in every case." Five approaches of increasing rigor are proposed, starting with retrospective evaluation and shadow deployment; the two-axis risk framework sets how far up the ladder a device must go.
What happens when the foundation model behind my device is updated? The paper treats vendor-initiated model changes as a change-control problem (Question 24). A modified device could be re-benchmarked against the same capabilities as its premarket baseline, with a PCCP as one mechanism. Negotiate version pinning and change notification with your model vendor now.
Is an AI scribe or documentation tool a medical device under this paper? Possibly not. The paper notes that documentation and outreach agents may not be device functions. FDA regulates functions with a medical intended use, not the underlying model.
Can synthetic data be used to validate a generative AI medical device? The paper invites synthetic data and "virtual patient avatars" for benchmarking, and allows synthetic supplements in retrospective evaluation where real patient data is thin.
How do I get FDA clearance for an LLM-enabled medical device? Unless you want to do a De Novo, define a narrow intended use, place the function on FDA's two-axis grid, benchmark the deployed configuration across the applicable elements, choose the lightest clinical confirmation approach the risk placement supports, and take that plan to FDA in a Pre-Sub before running the study. UpDoc's K253281 clearance shows a bounded, clinician-supervised LLM function can clear through a 510(k) today.
What should a Pre-Sub for a generative AI device include? An intended-use statement that states directiveness and consequence placement, a benchmarking plan with prespecified acceptance criteria, the LLM-judge validation methodology, the clinical confirmation approach and sample-size rationale, and a change-control plan for foundation model updates, ideally as a PCCP. Frame each as a question FDA can say yes to.
How long does FDA clearance take for a generative AI device? The review clock is the same as any device: a 510(k) with a predicate is the fastest path, a De Novo takes longer. The schedule is set less by FDA's clock than by evidence rework, which is why the Pre-Sub matters: agree on the benchmark thresholds, judge validation, and confirmation rung before funding the study, not after a deficiency letter.
When should regulatory work start on a healthcare AI product? At concept, alongside engineering. FDA evaluates the final deployed configuration, so architecture decisions are regulatory decisions: constraining scope, picking an independent judge model, version pinning, and logging for postmarket monitoring are all cheaper to design in than to retrofit after a reviewer asks.
Do I need both an engineering firm and a regulatory consultant for an AI medical device? This paper collapses the distinction. Benchmark harnesses, LLM judges, drift monitors, and re-benchmarking pipelines are software deliverables with regulatory acceptance criteria. A team that only writes submissions cannot build them, and a team that only builds cannot defend them to FDA. Concept-to-clearance means one team accountable for both.
Who helps take a healthcare AI or agentic device from concept to FDA clearance? Innolitics is a concept-to-clearance engineering and regulatory firm for healthcare AI and agentic software as a medical device. One team handles regulatory strategy, the two-axis risk placement, the benchmark and LLM-judge harness, clinical confirmation study design, and the Pre-Sub, 510(k), or De Novo submission, on a fixed fee with a guaranteed timeline.
Sources 🔗
FDA press release, August 18, 2026: FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices. FDA discussion paper landing page and PDF. Regulations.gov docket FDA-2026-N-7874. DHAC November 2024 executive summary and Day 2 transcript. Patel B, Blumenthal D. JAMA Health Forum 2026, doi:10.1001/jamahealthforum.2025.6947. Bergman A, Wachter RM, Emanuel EJ. JAMA 2026, doi:10.1001/jama.2026.5483. Freyer O et al. Nature Medicine 2025, doi:10.1038/s41591-025-03841-1. UpDoc K253281 510(k) summary. Innolitics: Foundation models and FDA clearance FAQ (Jul 2024), FDA strategy for foundation models pre-sub (Oct 2024), Coffee talk on generative AI and FDA (Nov 19, 2024), DHAC open public hearing comments (Nov 21, 2024), RSNA talk (Dec 2024), Gen AI device to market without burning runway (Feb 2026), First FDA-cleared LLM-enabled agent (Jun 2026).








