Public Comment on FDA's Generative AI Device Questions: A Running Census of Docket FDA-2026-N-7874

 September 22, 2026
SHARE ON

AI/MLRegulatory

Summary 🔗

FDA's August 2026 discussion paper on generative AI-enabled medical devices asks the public 26 questions. Two language models independently coded every public comment on docket FDA-2026-N-7874 against each question. As of September 21, 2026, the public has filed 107 comments, from 96 distinct submitters, and those comments contain 805 answers to individual questions. Nine of FDA's questions pose a proposition that a commenter can accept or reject. On those nine, 87% of positions accept FDA's proposition with requested changes, 5% accept it as written, 5% are mixed, and 2% reject it. The other 17 questions are open, and for those I report the recurring answers. The same requests appear throughout: account for reversibility of harm and speed of failure detection, evaluate the deployed system and not the model alone, keep a human in front of irreversible actions, and keep accountability with the manufacturer. In 67 comments both coders see a push toward more oversight; in seven, toward less. The figures update nightly until the docket closes on October 19, 2026.

About the author 🔗

Yujan Shrestha, Partner

Yujan Shrestha, MD

Partner

I have led or supported FDA strategy and submissions for 51+ medical devices, most AI-enabled, since co-founding Innolitics in 2012, across 510(k), De Novo, Breakthrough, Q-Sub, and deficiency responses. My work spans radiology AI (CADe, CADx, CADt, quantitative imaging), cardiac and neuro imaging software, dental AI, LLM and foundation model devices, and I have managed or contributed to 65 Innolitics projects. FDA's public clearance letters for AI Metrics (K202229, 2020) and Medweb's Lightning Viewer (K242362, 2024) name me as the sponsor's correspondent. I presented at FDA's first Digital Health Advisory Committee meeting on generative AI in November 2024, I have published 78 articles on innolitics.com, and I built fda.innolitics.com, our public FDA device search tool.

Connect on LinkedIn

Figure 1. Public response to each of FDA's 26 questions

Bar length is the number of distinct submitters who answer the question (n at right). Nine questions pose a proposition, and their bars are split by position. The other 17 are open questions and show one tone. Select a question to open it in Figure 2.

Submitter Position Summary of answer

Summaries are my paraphrase. Select a row to open the comment here, at the passage that supports my summary, highlighted; or follow the Regulations.gov link for the docket record. Attachments are copies of the files posted to the docket. Recurring answers were named by one model, and a comment counts toward one only when both models assigned it.

Risk assessment (Questions 1 to 6) 🔗

Of 52 submitters on the two-axis risk framework, 47 accept it with changes and two reject it. Twenty-four ask FDA to add reversibility of harm and time to harm as risk dimensions, 18 ask for traceability of outputs to source evidence, and 10 for the user's ability to detect an error. On directiveness (Question 2), eight submitters say it should be judged by functional effect and not wording, and eight say a disclaimer should not lower the classification. Question 3 draws the most divided response in the docket: of 26 submitters, 18 accept with conditions that patient-facing delivery carries different risk, six are mixed, and 10 caution FDA not to presume that it does. On multi-turn devices, 14 ask that whole conversations be the unit of evaluation. On care escalation, 10 ask that under-escalation and over-escalation be reported separately and never as one score.

Premarket evaluation (Questions 7 to 17) 🔗

All 37 submitters who take up the competency-based approach accept it, 31 of them with changes, and 11 ask that the final deployed system be evaluated and not the foundation model alone. On benchmark validity, 11 ask for sequestered test sets held by an independent party. On clinical confirmation, 13 would scale rigor to risk and autonomy, and nine would reserve prospective studies for high-consequence autonomous functions. On comparators, 18 ask for human-AI team evaluation where clinician review is intended, and 15 would allow comparators that reflect the care likely without the device. On synthetic data, nine see its place in stress tests and rare scenarios, and six want generators independent of the device's model family. On third parties, 13 see their role as holding sequestered datasets and adjudicating cases.

Postmarket monitoring and change (Questions 18 to 24) 🔗

No submitter accepts the trade of premarket evidence for postmarket monitoring without conditions: 30 of 32 accept it conditionally and two reject it. Fourteen would exclude autonomous, high-consequence, or irreversible functions, and nine require a prespecified monitoring plan with thresholds and triggers. On postmarket evaluation, 13 want a periodic floor combined with event-triggered reassessment, and 13 name model, prompt, retrieval, or tool changes as triggers. On stakeholder roles, 26 of 34 say the manufacturer retains accountability. On supervisory agents, 10 require independence from the device's own model. On third-party model changes, 18 ask for contractual advance notice and 15 for pinned versions with no silent updates. On change control plans, 12 would prespecify a change envelope and its boundaries instead of individual changes.

Other topics (Questions 25 and 26) 🔗

On foundation model master files, 25 of 27 accept them with changes; 12 specify content such as versioning, limitations, and update policy, and 10 say a master file must not shift responsibility from the sponsor. On agentic devices, 22 of 42 submitters ask for human authorization before irreversible or high-consequence actions, 13 for halt and rollback states, and 10 for tamper-evident logs that reconstruct what the agent did.

My position 🔗

I have argued since the November 2024 Digital Health Advisory Committee meeting that FDA should accept lighter premarket evidence for generative AI devices in exchange for strong postmarket evidence (my remarks are here). Premarket clinical evidence is the largest source of delay for the sponsors I work with, and it is a capital expense paid before a product earns anything. Postmarket surveillance is an operating cost that a prudent manufacturer carries anyway once version 1.0 ships.

The submitters are right that the trade depends on detection latency: if the first observable sign of failure is patient harm, monitoring is a record of injury and not a control. I disagree with excluding high-risk functions as a category. Risk category is a proxy for reversibility and detection latency, and some high-risk functions score well on both. I also agree with the submitters who would define a change control plan by its envelope. A sponsor cannot predict what a foundation model vendor will change, but it can fix the tests every change must pass.

Limitations 🔗

The sample is self-selected. Trade associations, foundation model developers, and health systems have not yet filed. The coding was done by language models under rules I set, and I have not read every comment myself. The count of comments pushing toward more or less oversight is a separate whole-comment judgment on which the models agreed less (kappa 0.70), so it includes only comments where both agreed.

Materials and methods 🔗

Source material. I retrieved every comment on the docket through the Regulations.gov API, with its 92 PDF and Word attachments, and extracted the text: 107 comments and about 290,000 words. One submission is a verbatim copy of FDA's own paper and is coded as answering nothing. Submitters who filed more than once are counted once, which gives 96 submitters.

Unit of analysis. Each of FDA's 26 questions is coded separately. Nine questions pose a yes-or-no proposition (Questions 1, 3, 7, 9, 16, 17, 18, 20 and 25), and for those a coder records a position: accepts, accepts with changes, mixed, or rejects. The other 17 ask how or what, and for those a coder records only whether the comment answers the question. Figure 2 shows FDA's wording and, where there is one, the proposition coded.

Coding panel. Claude Fable 5.1 and GPT-6 Astra each coded every comment independently from the same written rubric. For each question a coder recorded whether the comment answers it, a position where applicable, a verbatim passage, and up to four key points. Passages were string-matched against the source and discarded if inexact. A comment answers a question when it cites the question or makes a concrete point that lands on its subject; a one-sentence remark in a comment scoped to other questions does not count.

Reconciliation. The models agreed on 96.8% of 2,782 coding decisions (107 comments times 26 questions; Cohen's kappa 0.93). Where they differed, each re-read the comment alongside both anonymised codings and decided again, which settled 67 decisions. I adjudicated the remaining 25 myself.

Adjudication rules. An earlier pass grouped the questions into eleven topics. I adjudicated its 25 split decisions under four rules, which were then written into the rubric for the per-question pass:

  1. A comment that endorses a proposition and then lists safeguards or implementation details accepts it with changes.
  2. A comment that proposes a larger model containing FDA's two risk axes accepts the framework with changes. One such comment calls the two axes "a good starting point."
  3. A one-sentence endorsement inside a comment scoped to other questions does not count as an answer.
  4. A concrete request that lands on a question's subject counts as an answer even when the comment does not cite the question.

The 21 split decisions in the per-question pass fell into three patterns. In six, both models agreed the question was answered, but one mislabeled a proposition question as open, and I took the other model's position. In 11, the models split on whether a point belonged to this question or a neighboring one, and I left it out of this question. Four repeated earlier rulings.

Recurring answers. One model named four to seven recurring answers per question from the coded material, and that list is frozen. A comment counts toward an answer only when both models assign it, so those counts are floors.

I analyzed the paper itself in an earlier article.

Results 🔗

Figure 1 shows how many submitters answer each question. The most answered are Question 19 on postmarket evaluation (56 submitters), Question 1 on the risk framework (52), Question 9 on the benchmarking structure (43), and Question 26 on agentic devices (42). The least answered is Question 17 on other model architectures (10). Selecting a question in Figure 1 opens it in Figure 2, which shows FDA's wording, the recurring answers, and the comments behind them.

Updates 🔗

The figures are regenerated nightly. I will revise the text if the distribution changes materially, and once after the docket closes.

Further reading 🔗

SHARE ON

From concept to FDA clearance. One partner.

Software, AI/ML, cybersecurity, clinical validation, QMS, and the submission itself under one roof. 50+ FDA-cleared SaMD, software documentation in four months, 510(k) clearances in as little as nine months, backed by our timeline and clearance guarantee.

We Got a Cardiac MRI Analysis AI SaMD Cleared in 5.5 Months, PCCP Included

First submission, first clearance: multi-model cardiac MRI AI cleared in 5.5 months with an FDA-authorized PCCP.

Heartvue.ai - Dr. Jeffrey Dendy logo

We Mapped the Fastest Path to FDA Clearance for AI/ML SaMD

Existing EU data. Strategy designed to get to U.S. market with minimal extra effort.

AUMI ai logo

We went from idea to FDA clearance in 18 months and first Install for AI Tumor Tracking Software

Prototype → 510(k) → QMS → FDA audit passed. Multi-year partnership with ongoing development.

AI Metrics - Dr. Andrew Smith logo

We Did It All: R&D, Clinical Study, and 510(k) Cleared in 12 Months

Built and cleared an AI cardiomegaly detector end-to-end: algorithm R&D to 510(k) clearance in under 12 months.

Body Check logo

We helped a University Spinout go from Prototype to Acquisition.

We secured breakthrough status and drove a 3-month FDA submission for an AI breast-risk SaMD.

University spinout acquired by global AI company logo

We Submitted a 510(k) in 7 Weeks and Cleared it in 4 Months

Completed DHF and 510(k) documentation in 7 weeks. FDA cleared 4 months after submission."

Medweb logo

We Secured a Breakthrough Device Designation for a ECG Foundation Model

Reframed regulatory strategy after FDA pushback. BDD granted in ~3 months.

Trusted by

  • nvidia
  • University of Alabama
  • Mary Bird Perkins
  • PhotoniCare
  • Butterfly Network
  • Enlitic
  • OXOS
  • NSI
  • Prenuvo
  • Transonic
  • AI Metrics
  • RadUnity
  • Varian-Mobius
  • Echo IQ
  • Envisionit
  • Smile Dx
  • Indica Labs
  • Magnetic Insight
  • Neosoma
  • BodyCheck

Learn more about our AI/ML 510(k) Submission service.

Our FDA Clearances

Related Articles

Let's Talk

Every great partnership starts with a conversation. Fill out the form below for a discovery call, and an Innolitics team member will contact you soon.

Book a 30-minute call →