---
author:
- Yujan Shrestha
cta_button_text: Take my GenAI device to clearance
cta_case_studies:
- 3c9bd5b7-a754-81db-9083-f5aca6838d83
- a4b2d790-f1ff-47bb-93b2-09b0d4619030
cta_hero_details: Innolitics builds the software, trains the AI, writes the risk file
  and postmarket plan, and prepares the submission under one roof, including GenAI
  devices with HCP on the loop. 50+ FDA-cleared SaMD, software documentation in four
  months, and 510(k) clearances in as little as nine months, backed by our timeline
  and clearance guarantee.
cta_hero_text: From concept to FDA clearance. One partner.
date: '2026-10-10'
description: Yujan Shrestha, J. David Giese, and former FDA advisory committee lead
  James Swink discuss FDA\'s generative AI discussion paper, a calibrated risk grid,
  patient-facing devices, and lighter premarket proof in exchange for stronger postmarket
  monitoring. The docket closes October 19, 2026.
related:
- 3e1bd5b7-a754-815a-9958-d37826744584
- 3c0bd5b7-a754-81ba-be69-db85b30e7ecf
- 147bd5b7-a754-8068-90c1-f1406ae43b8c
- 3cfbd5b7-a754-80c6-ad6d-c222fe90f1a7
- 38abd5b7-a754-814c-b39a-d204c149640f
- 3e3bd5b7-a754-80e1-9bf9-faa1471aba9c
- 3b6bd5b7-a754-8084-bd41-ce559b39aabb
- b78c05f0-1406-4bf8-aa5b-23410b0cbfe1
- e0f66593-9702-42ec-9983-3da8d19ce690
- 307bd5b7-a754-801b-9977-fce85266969e
- 3f5bd5b7-a754-814e-ae81-d3be19dd5566
- 3f5bd5b7-a754-81ca-b731-f821d765edfb
- 3f5bd5b7-a754-81ae-bdb9-f51de35db8f4
- 3f5bd5b7-a754-819c-98fe-f745b239827e
- 3f5bd5b7-a754-81b0-b72d-d237cb06fffa
- 3f5bd5b7-a754-811b-8773-df7b2c68022b
- 3f5bd5b7-a754-8189-a0c7-f5182e9f56c4
- 3f5bd5b7-a754-813d-a174-eb85b08a7f29
- 3f5bd5b7-a754-8103-ac69-e0711762bed6
- 3f5bd5b7-a754-81fc-afb8-c5980d840a19
- 3f5bd5b7-a754-81ad-bdda-d7e8d1f8dbe2
- 3f5bd5b7-a754-81e8-b9ad-de4b2b581393
- 3f5bd5b7-a754-81d3-a350-ece68991a979
- 3f5bd5b7-a754-8170-b7a2-d583570ef103
- 3f5bd5b7-a754-8101-8df4-f139f3de116e
title: 'FDA''s Generative AI Discussion Paper: Notes from a Coffee and Q&A on Risk
  Tiers, Patient-Facing Devices, and the Premarket-Postmarket Trade'
topics:
- Regulatory
- AI/ML
---

### Summary

I want FDA to accept lighter premarket proof when a manufacturer commits to stronger postmarket monitoring. That was my main argument when James Swink hosted J. David Giese and me on LinkedIn Live on October 8, 2026, to discuss FDA\'s August 2026 paper on generative AI-enabled medical devices and its docket comments. David agreed, though he pointed to FDA staffing and patient privacy as obstacles. James focused on enforcement and said FDA would need triggers and a way to act when a manufacturer does not follow its postmarket plan.

**Key points**

- The premarket-postmarket trade benefits FDA and industry. As of October 8, 2026, **39 of 42** census submitters on the question accept it with conditions.
- My risk grid places existing FDA determinations in FDA\'s cells. The acute-harm HCP-on-the-loop column has no FDA determination to anchor it.
- Clinician review after a device acts controls risk only while the harm can still be reversed.
- Small workflow changes can move one stroke product from non-device CDS to a PMA.
- Patient-facing risk can be measured only through postmarket telemetry.
- David and I each reject a central board exam for current GenAI devices.

<p class="relative pb-[56.25%] pt-[25px] h-0 w-full"><iframe width="560" height="315" src="https://www.youtube.com/embed/ADra_F3iU14" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" allowfullscreen class="absolute top-0 left-0 w-full h-full"></iframe></p>

*Full recording of the October 8, 2026 Coffee and Q&A (64 minutes, edited for pauses, with captions and chapters).*

### Background

FDA published *Considerations for the Regulation of Generative AI-Enabled Medical Devices* on August 18, 2026, under docket FDA-2026-N-7874 \[1, 2\]. The paper\'s 26 questions cover risk, premarket evaluation, postmarket monitoring, and foundation models and AI agents. It is neither draft nor final guidance. James described it as FDA thinking out loud and asking for help. An earlier Innolitics article summarizes the paper \[12\].

James Swink joined Innolitics as Principal Regulatory Specialist in August 2026. He spent 22 years at FDA\'s Center for Devices and Radiological Health and directed more than 150 advisory committee meetings. J. David Giese and I founded Innolitics in 2012 and are Partners.

<figure>
  <img src="/img/articles/FDAs_Generative_AI_Discussion_Paper_Notes_from_a_Coffee_and_QA_on_Risk_Tiers_Patient-Facing_Devices_and_the_Premarket-Postmarket_Trade-f21ea738902c4883a137bcce97c7bbc2.png">
  <figcaption>
    Figure 1. The three participants in the October 8, 2026 Coffee and
    Q&amp;A session: James Swink (host), J. David Giese, and Yujan Shrestha.
  </figcaption>
</figure>

James opened with a person asking a health app about a tight chest at 2 AM. The app answers in seconds, confident and warm, advising whether to go to the emergency room. James asked who had checked that answer. He also reminded listeners of the risk at the other end.

<div class="pl-3 border-solid border-0 border-l-4 border-gray" custom-style="Block Quote" markdown="1" role="blockquote">

\"A safe device that never reaches patients does not help anyone either.\" (James Swink, quoted from the session)

</div>

### Generative AI and the limits of labeling

For me, the difficulty of regulating GenAI devices starts with labeling. Every submission rests on it. FDA began as a guard against mislabeled products, and the label still underpins its regulation. A reviewer reads the claim, asks for scientific proof, and authorizes the label.

That approach works when the label stays fixed, as it did for the narrow algorithms before GenAI. A generative model can tell my toddler a joke or help assess whether I might have cancer. The intended use changes with the prompt.

David ranked this as the biggest problem and added two others. In his view, GenAI is the first technology that behaves intelligently. Its intended use is broad and hard to bound, and manufacturers have trouble preventing patients from changing it.

Third-party model changes were his second concern, which he ranked below the first. OpenAI, Anthropic, and Google do not currently guarantee a frozen version of the models that many devices run on. His third concern was interactivity. A mental health app such as Limbic talks with a patient over many turns and can complete tasks on its own. David said testing that behavior differs from testing the more than 1,600 traditional AI-enabled devices on FDA\'s list.

### Public comments on FDA\'s questions

I run a nightly census of the docket\'s public comments \[4\]. Two language models, one from Anthropic\'s Claude family and one from OpenAI\'s GPT family, use the same prompt to code each comment against FDA\'s 26 questions. I use two vendors for this LLM-as-a-jury design because their models have different biases. When they disagree, each model re-reads the comment alongside both codings and decides again. I adjudicate the remaining disagreements under written rules.

James read figures as of September 21, 2026, during the session. The figures here are as of October 8, 2026. The **129** comments from **115** submitters contain about **390,000** words. Across **3,354** coding decisions, the models agreed on **96.8%** (Cohen\'s kappa 0.93). Re-reading settled 80 more, and I adjudicated 23. Six decisions remain disputed. I exclude those from all counts.

Nine of FDA\'s questions pose a proposition that submitters can accept or reject. Of 370 positions on those nine, **88%** accept FDA\'s proposition with requested changes, **5%** accept it as written, **5%** are mixed, and **2%** reject it. Both models see a push toward more evidence or oversight in **81** comments and toward less in **10**.

Most answers amount to \"yes, but.\" The conditions tell me what each submitter cares about.

<figure>
  <img src="/img/articles/FDAs_Generative_AI_Discussion_Paper_Notes_from_a_Coffee_and_QA_on_Risk_Tiers_Patient-Facing_Devices_and_the_Premarket-Postmarket_Trade-4890b376977b4335a0dbd2308554234a.png">
  <figcaption>
    Figure 2. Census of submitter positions on FDA's 26 questions, docket
    FDA-2026-N-7874, as of October 8, 2026. The census updates nightly until
    the docket closes.
  </figcaption>
</figure>

David cautioned against treating the docket as an unbiased sample of U.S. opinion. Many submitters work in medical device regulation. As of October 8, 2026, companies, regulatory professionals, consultancies, law firms, and trade associations filed 58 of the 129 comments (45%). In his view, most patients are glad to have ChatGPT, and new FDA controls on it would provoke a political outcry. He called comments from people with clinical or engineering experience extremely valuable.

I asked patients, caregivers, and the public to comment because the public\'s appetite for risk sets the benefit-risk calibration point. Society can choose to move faster and accept some risk, or demand more premarket proof and get fewer devices of higher quality.

### A risk grid anchored to FDA precedent

FDA\'s proposed risk grid crosses the consequence of an incorrect output with the function\'s activity, from informational and non-directive to fully autonomous action \[2\]. Where a device lands would tell a manufacturer how much evidence FDA expects. David thought two axes could not capture the range of GenAI devices. He preferred a full ISO 14971 risk assessment, while still seeing value in a heuristic for the level of evidence.

I think consequence and activity explain the most variance. Other dimensions, including time to harm, still matter. Guidance has to serve the whole industry and stay current for years, which leaves little room for specifics. I went one step further and placed calibration points in FDA\'s cells, each drawn from an existing FDA determination.

<figure>
  <img src="/img/articles/FDAs_Generative_AI_Discussion_Paper_Notes_from_a_Coffee_and_QA_on_Risk_Tiers_Patient-Facing_Devices_and_the_Premarket-Postmarket_Trade-a4b3c7a7dc1b4d06a0760af6006fd8b8.png">
  <figcaption>
    Figure 3. A hypothetical risk matrix for GenAI devices with five
    consequence rows and seven activity columns. Suggested levels are 1
    outside FDA oversight, 2 enforcement discretion, 3 Class II, and 4 Class
    III PMA. A crosswalk maps the columns to FDA and IMDRF N81 terms.
  </figcaption>
</figure>

- **Patient-facing, informational, non-directive.** A person uses a chatbot like a search engine. I expect this use to stay unregulated, as Google has.
- **Patient-facing, action-directing.** Natural Cycles (DEN170052), which predicts fertile days for contraception, anchors the serious row at Class II.
- **Negligible consequence (bottom row).** General wellness advice, such as a daily walk, carries too little risk to quantify.
- **HCP in the loop, reviewable basis.** A clinician can verify the basis of the output. FDA\'s CDS guidance governs this column, and most functions here are non-device CDS \[5\].
- **HCP in the loop, non-reviewable basis.** Most AI software FDA clears today sits here, including CADe, CADt, and ECG analysis devices.
- **HCP out of the loop.** The automated external defibrillator acts autonomously, with a catastrophic consequence of error, and requires a PMA. IDx-DR (DEN180001) detects diabetic retinopathy without a clinician and anchors the serious row at Class II.
- **HCP on the loop.** FDA calls this \"action-taking, HCP-supervised.\" I split it into slow-harm and acute-harm columns and said on air that little precedent exists for either. My written comment places Class II determinations such as UpDoc in the slow-harm column. It suggests enforcement discretion, a lighter level, for slow harm at minor or serious consequence. No FDA determination anchors the acute-harm column.

Before GenAI, there was little reason for an on-the-loop workflow to exist. David first asked how I distinguish \"in the loop\" from \"on the loop.\" If a device escalates to a clinician and blocks until the clinician responds, the HCP stays in the loop. If it escalates while also telling the patient what to do, the HCP is on the loop. I treated the three rightmost columns as action-taking because the AI acts before a clinician decides.

David then asked for an example in the lower consequence rows, which I had not yet calibrated. He suggested a mental health app such as Limbic, which is not indicated for suicidal patients. I agreed. Active suicidal ideation would put the same device in the acute-harm column, at serious consequence or higher.

### Time to irreversible harm

Putting an HCP on the loop controls risk only if the clinician can correct a wrong action before the harm becomes irreversible. Once the action cannot be reversed, later review controls nothing. **31** census submitters ask FDA to add reversibility of harm and time to harm as risk dimensions.

The time to irreversible harm also determines the review deadline. Suppose a chatbot schedules a biopsy for two weeks from now. The risk file can require HCP review within those two weeks. James pointed out that some drugs allow a much shorter window for reversal. Psychological harm to the patient counts too.

### Traceability and the reviewable basis

For the reviewable-basis columns to work, a clinician must be able to trace an output to its sources. **23** submitters ask for traceability of outputs to source evidence. David said commenters use traceability to mean both citations and an account of how a model arrived at its output. The varied inputs that make these devices useful also make traceability impossible in some cases.

<div class="pl-3 border-solid border-0 border-l-4 border-gray" custom-style="Block Quote" markdown="1" role="blockquote">

\"In the cases where you always can have traceability, an LLM is less necessary anyway.\" (J. David Giese, quoted from the session)

</div>

UpDoc separates the roles by putting an LLM conversation around a deterministic core \[6\]. David considers traceability infeasible for mental health, post-surgery follow-up, and weight counseling. He recommends it where it is feasible.

### One stroke product, three regulatory paths

James adapted a stroke example from FDA\'s paper into a hypothetical startup. Its AI reads a stroke scan and orders a clot-busting drug without waiting for a neurologist. My first point to that company would be how much a subtle change in intended use or workflow can change its place in the grid.

1.  An LLM applies clinical guidelines to the radiologist\'s report and lists drug options for another clinician. The HCP is in the loop and can review the basis. This might be non-device CDS, although the serious and time-critical setting may disqualify it.
2.  A multimodal model analyzes the images and recommends the drug. The clinician cannot easily check how the model read the image, so the basis is non-reviewable. Most devices in this column are Class II.
3.  A patient with stroke symptoms walks into a CVS pharmacy. Pharmacy staff act on the device\'s output and send the case to a clinician for review. The HCP is on the loop, with acute harm.

Each step raises the risk and likely the device class. The cost of a 510(k), a De Novo, or a PMA can range from 1x to 100x. I ask which variation fits the company\'s business model and customers. That is the variation that gives it the highest return on its regulatory spend.

### Patient-facing devices and paternalism

James introduced patient-facing devices as one of the docket\'s most divided questions. He asked where protecting patients turns into paternalism.

FDA\'s lisinopril example gives four phrasings, starting with a general statement about dose increases and ending with a direct instruction to raise the dose from 10 mg to 20 mg daily. James asked whether adding \"but check with your doctor first\" changes the risk. Labeling is a risk control. But when the message directs a patient to act, I doubt that sentence changes the type of device. David said labeling is even less useful for patient-facing devices because most patients do not understand it. I agreed that few patients read a drug label in detail.

On FDA\'s Question 3, **12 of 31** submitters caution FDA against presuming that patient-facing functions carry higher risk. David noted that many submitters are clinicians. They have a subtle conflict of interest, he said, because they tend to believe patients need clinicians. Patients have time and a direct stake in the outcome.

James observed that someone with a lifelong condition may know more about it than some family practitioners. David agreed, drawing on experience with family and friends. He sees different risks in AI that talks to patients and AI that talks to doctors, and he favors enabling patients.

In my view, no one can know this risk in the premarket. Many people in the United States already use ChatGPT for health questions, including my parents. If that use were very dangerous, I would expect to hear about many more adverse events, even without a duty to report them. There is still a large public exposure that has not been measured.

Medical school trains a physician for the 0.1% of patients who need a second question, such as the patient with a suspicious mole. For the other 99%, I expect the benefits to outweigh the risks. To find the exceptions, these devices need to reach the market. Then manufacturers can measure telemetry and adverse events and carve out the cases where risks exceed benefits. Until there are postmarket data, every estimate of this risk is hypothetical.

### The board-exam analogy for premarket evidence

The hardest premarket case is a device that lets a generalist work like a specialist. Point-of-care ultrasound guidance already does this. GenAI raises the problem of how to test that claim before launch. On air, I asked whether anyone would make a family physician take a cardiology board exam with ChatGPT.

FDA also asks whether a device could be evaluated through a board exam and supervised practice, much as a physician is credentialed. David answered \"no.\" The GenAI devices Innolitics clients build are narrow. Validating them looks closer to traditional AI validation than to a medical school exam. He sees a possible fit only if FDA regulated ChatGPT itself.

In my experience, the model that leads a general-purpose coding or reasoning benchmark often performs worse on my use cases than another model. Medicine is more complex than those tasks. No one yet understands the failure modes well enough to design a board exam that would mean anything.

<div class="pl-3 border-solid border-0 border-l-4 border-gray" custom-style="Block Quote" markdown="1" role="blockquote">

\"If you\'re doing a med device, you need to own your own benchmark.\" (Yujan Shrestha, quoted from the session)

</div>

With enough data, a central benchmark for a specific failure mode may make sense. For now, I keep coming back to premarket clearance with postmarket telemetry built into it, monitored through automated and manual methods. I expect industry to find that palatable. It also provides real evidence. FDA\'s Question 18 asks about the same arrangement.

### Lighter premarket proof for stronger postmarket monitoring

Question 18 asks whether FDA could accept more uncertainty before launch if a manufacturer commits to close monitoring afterward \[2\]. FDA\'s answer sets how much evidence a manufacturer must fund before it earns revenue.

As of October 8, 2026, **39 of 42** submitters on this question accept the trade with conditions, **3** reject it, and none accept it without conditions. **20** would exclude autonomous, high-consequence, or irreversible functions. **10** require a prespecified monitoring plan with thresholds and triggers. James pointed out that I made the same argument at FDA\'s November 2024 Digital Health Advisory Committee meeting on generative AI, where he served as FDA\'s Designated Federal Officer \[3\].

Most decisions I help clients make involve a trade-off. This is a rare one where both parties gain.

<div class="pl-3 border-solid border-0 border-l-4 border-gray" custom-style="Block Quote" markdown="1" role="blockquote">

\"I think this is one of these things where it\'s a win-win for everybody involved. Win-win for FDA, win-win for industry.\" (Yujan Shrestha, quoted from the session)

</div>

#### Benefit to FDA

FDA wants evidence that a device generalizes to real patients. In a premarket submission, a manufacturer presents a controlled study and argues that the results will hold after launch. No one can prove that in advance.

Postmarket data show generalization directly. They come from the intended-use population, which makes them, in my view, the highest grade of evidence. FDA pays for that evidence by accepting more uncertainty at clearance.

#### Benefit to industry

A manufacturer pays for premarket evidence before it has revenue, using money raised at a high cost of capital. That is a capital expense. Postmarket monitoring can be a cost of goods sold, paid from revenue. Most business owners will easily take a cost of goods sold over a capital expense. A manufacturer with a good product monitors it anyway.

#### Cost of waiting

It is difficult to quantify the benefits and risks of GenAI devices before launch. I am concerned about getting stuck in a long loop of analysis paralysis. A lightweight device on the market produces real outcome data. Even general controls would make these devices safer.

#### Existing mechanisms

A De Novo grant can include special controls requiring postmarket data within a set period. Section 522 lets FDA order postmarket surveillance studies \[7\]. I also suggested that a predetermined change control plan (PCCP) could carry the monitoring commitment and serve as its enforcement mechanism, because changes outside the plan require a new submission \[8\].

David agreed and cited the TEMPO pilot. FDA uses enforcement discretion there to let a device reach the market and collect real-world data, with a 510(k) later \[9, 10\]. In his article on Limbic, David describes the company starting with lower-risk indications and widening them as real-world data calibrate the risk \[11\].

#### Staffing, data access, and enforcement

David saw two obstacles. FDA allocates staff and budget around 510(k) review, so routine postmarket review could require reorganization. Patient privacy and hospital data-sharing agreements can also restrict access to postmarket data. One of David\'s current clients faces that problem.

James answered the staffing concern.

<div class="pl-3 border-solid border-0 border-l-4 border-gray" custom-style="Block Quote" markdown="1" role="blockquote">

\"It\'s the same review, it\'s just different data.\" (James Swink, quoted from the session)

</div>

The same reviewers would assess reports at 3, 6, 9, and 12 months. In James\'s view, enforcement is FDA\'s real concern. The agency needs a trigger and a way to remove a device from the market when a company fails to follow its postmarket plan.

He said the plan must contain floors, triggers, retesting, near misses, overrides, subgroup results, and real-world data. Ideally, a reviewer could watch weekly data under the 510(k) number and call the company when a trigger fires. James recommended that industry press FDA on low-risk devices first.

### Recommendations for teams building GenAI devices

- **Test-driven development.** If a GenAI product may be regulated in the future, write strong benchmarks and evals for each clinical task before writing the prompts that perform it. Add risk controls. I consider this good GenAI engineering practice anyway, and the work pays off if FDA regulates the product.
- **Docket comment.** Comment on FDA-2026-N-7874 before October 19, 2026. Patients and members of the public filed 11 of 129 comments.

I ended with advice I introduced on air as a familiar saying.

<div class="pl-3 border-solid border-0 border-l-4 border-gray" custom-style="Block Quote" markdown="1" role="blockquote">

\"Don\'t build devices that you wouldn\'t use on your grandma.\" (Yujan Shrestha, quoted from the session)

</div>

James closed by pointing listeners to the nightly census \[4\] and FDA Device Explorer, a free predicate search tool Innolitics built \[13\].

### Limitations

This article reports a 68-minute conversation. It is not a study, and the positions belong to the speakers. The census figures are as of October 8, 2026, and will change before the docket closes. I removed filler words from the quotations while keeping their meaning.

### References

1.  FDA. *Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback.* August 18, 2026. <https://www.fda.gov/medical-devices/digital-health-center-excellence/considerations-regulation-generative-ai-enabled-medical-devices-discussion-paper-and-request>
2.  FDA. Discussion paper PDF. <https://www.fda.gov/media/194242/download>. Docket FDA-2026-N-7874: <https://www.regulations.gov/docket/FDA-2026-N-7874>
3.  Shrestha Y. Yujan\'s Comments at FDA Digital Health Advisory Committee on GenAI Enabled Devices. Innolitics, November 2024. <https://innolitics.com/articles/yujan-fda-gen-ai-ac-meeting/>
4.  Shrestha Y. Public Comment on FDA\'s Generative AI Device Questions: A Running Census of Docket FDA-2026-N-7874. Innolitics, 2026. <https://innolitics.com/articles/fda-generative-ai-device-public-comments-census/>
5.  FDA. *Clinical Decision Support Software.* Guidance for Industry and FDA Staff. <https://www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software>
6.  Innolitics. First FDA-Cleared AI Agent and LLM Enabled Device Confirmed. <https://innolitics.com/articles/updoc-fda-cleared-ai-agent/>
7.  FDA. 522 Postmarket Surveillance Studies Program. <https://www.fda.gov/medical-devices/postmarket-requirements-devices/522-postmarket-surveillance-studies-program>
8.  Innolitics. PCCPs: Best Practices, FAQs, and Examples. <https://innolitics.com/articles/pccps-best-practices-faqs-and-examples/>
9.  FDA. TEMPO Digital Health Devices Pilot. <https://www.fda.gov/medical-devices/digital-health-center-excellence/tempo-digital-health-devices-pilot>
10. Innolitics. Cadence Enters TEMPO: What the Second Participant Suggests About the Pilot. <https://innolitics.com/articles/cadence-tempo-treatment-directing-software/>
11. Giese JD. Early Lessons on What It Takes to Get a Generative AI-Enabled Medical Device. Innolitics, September 2026. <https://innolitics.com/articles/generative-ai-study-design/>
12. Innolitics. FDA Releases Generative AI Medical Device Discussion Paper: What\'s Inside. <https://innolitics.com/articles/fda-generative-ai-medical-devices-discussion-paper/>
13. FDA Device Explorer, a free 510(k), De Novo, and predicate search tool. <https://fda.innolitics.com/>

### Further reading





### About the author


