---
author:
- Yujan Shrestha
cta_button_text: Take my concept to clearance
cta_case_studies:
- 3c9bd5b7-a754-81db-9083-f5aca6838d83
- 5a6365b6-8971-407c-9e10-f60e6a194d08
- 35d69707-1cbd-4b24-b365-f8bdbe34ab6d
- 141bd5b7-a754-80c4-8e9a-fa24807a2e59
- 35cbd5b7-a754-80f4-9a3c-cc0e62e41d52
- 170bd5b7-a754-8008-92e4-f68616c03660
- a4b2d790-f1ff-47bb-93b2-09b0d4619030
cta_hero_details: Software, AI/ML, cybersecurity, clinical validation, QMS, and the
  submission itself under one roof. 50+ FDA-cleared SaMD, software documentation in
  four months, 510(k) clearances in as little as nine months, backed by our timeline
  and clearance guarantee.
cta_hero_text: From concept to FDA clearance. One partner.
date: '2026-09-21'
description: 'A nightly-updated census of every public comment on FDA\''s generative
  AI medical device discussion paper (docket FDA-2026-N-7874), coded against all eleven
  proposals by a two-model panel. 88% of positions support FDA\''s direction with
  requested changes, 2% oppose it, and the same requests recur: reversibility of harm,
  detectability of failure, and sponsor accountability.'
related:
- 3c0bd5b7-a754-81ba-be69-db85b30e7ecf
- 3cfbd5b7-a754-80c6-ad6d-c222fe90f1a7
- b78c05f0-1406-4bf8-aa5b-23410b0cbfe1
- 38abd5b7-a754-814c-b39a-d204c149640f
- 353bd5b7-a754-8055-8506-f19a9e8f91a1
- 121bd5b7-a754-8024-a8ec-dab9ea61932e
- 4f4d00e9-04e8-428d-9a65-2b273979ef30
- 130bd5b7-a754-80c2-adae-e223512be4ac
title: 'Public Comment on FDA''s Generative AI Device Proposals: A Running Census
  of Docket FDA-2026-N-7874'
topics:
- Regulatory
- AI/ML
---

### Summary

FDA\'s August 2026 discussion paper on generative AI-enabled medical devices asks the public 26 questions. I grouped them by subject into eleven proposals, from the two-axis risk framework to foundation model master files, and two language models independently coded every public comment on docket FDA-2026-N-7874 against each one. As of September 20, 2026, the public has filed 98 comments, from 87 distinct submitters. A comment can take up any number of the eleven proposals; counted one submitter and one proposal at a time, the docket holds 411 positions. Of those positions, 88% support FDA\'s direction with requested changes, 7% support it without changes, and 2% oppose it. The same requests recur: account for reversibility of harm and speed of failure detection, evaluate the deployed system and not the model alone, and keep accountability with the sponsor. In 59 comments both coders see a push toward more oversight; in seven, toward less. The figures update nightly until the docket closes on October 19, 2026.

### Materials and methods

**Source material.** I retrieved every comment on docket FDA-2026-N-7874 through the [Regulations.gov](http://regulations.gov/) API, with its 84 PDF and Word attachments, and extracted the text: 98 comments and about 280,000 words. One submission is a verbatim copy of FDA\'s own paper and is coded as taking no position. Submitters who filed more than once are counted once, which gives 87 submitters.

**Coding panel.** Claude Fable 5.1 and GPT-6 Astra each coded every comment independently from the same written rubric. The eleven proposals are my grouping of FDA\'s 26 questions; Figure 1 lists the question numbers under each. Question 26, on agentic devices, is not covered by any of the eleven. For each of the eleven proposals a coder recorded whether the comment addresses it, a position (support, support with changes, mixed, oppose), a verbatim passage, and the specific changes requested. Passages were string-matched against the source and discarded if inexact. A \"yes, provided\" answer and a \"no, unless\" answer that state equivalent conditions both count as support with changes.

**Reconciliation.** The models agreed on 94.1% of 1,078 coding decisions (98 comments times eleven proposals) (Cohen\'s kappa 0.89), and on 97.3% of whether a proposal was addressed at all. Where they differed, each re-read the comment alongside both anonymised codings and decided again, which settled 39 decisions. I adjudicated the remaining 25 (2.3%) myself.

**Adjudication rules.** The 25 split decisions fell into four patterns, and I decided each pattern once.

1.  A comment that endorses a proposal and then lists safeguards or implementation details is support with changes, not support (6 decisions). Example: endorsing third-party testing subject to independence and conflict-of-interest criteria.
2.  A comment that proposes a larger model containing FDA\'s two risk axes is support with changes, not opposition (2 decisions). One such comment calls the two axes \"a good starting point\" and adds model and lifecycle layers.
3.  A one-sentence endorsement inside a comment scoped to other questions does not count as addressing the proposal (3 decisions).
4.  A concrete request that lands on a proposal\'s subject counts as support with changes even when the comment never names the proposal (11 decisions). Example: asking for multi-turn evaluation without mentioning patient-facing devices. This is the most consequential rule, and a stricter reader would drop these 11.

Three decisions fit no pattern: a comment that favors third-party testing but doubts it is feasible is mixed; a hedged sentence in a comment that scopes third-party testing out is not addressed; and a recommendation to use synthetic test suites for drift detection counts as a requested change on synthetic data.

**Recurring requests.** One model named four to seven recurring requests per proposal from the coded material, and that list is frozen. A comment counts toward a request only when both models assign it, so those counts are floors.

I analyzed the paper itself in [an earlier article](https://innolitics.com/articles/fda-generative-ai-medical-devices-discussion-paper/).

### Results

Figure 1 shows the positions on each proposal. Support without requested changes is rare. It is highest for clinical confirmation (6 of 41 submitters) and the generalist versus specialist question (4 of 16), and absent for synthetic data (0 of 23) and the premarket-for-postmarket trade (0 of 30). Selecting a proposal in Figure 1 opens it in Figure 2, which lists the recurring requests and links to the comments behind them.

<!-- GenAI docket comment census embed. Generated by research/genai-docket-comments-2026-09/build_embed.py. To update: replace this Notion HTML embed with the regenerated file. -->
<style>
#genai-census{--navy:#2D3F86;--navy2:#8F9BD0;--navy3:#E6E9F6;--orange:#E85036;--ink:#161B33;--ink2:#4A5172;--ink3:#7B819D;--rule:#D9DCEA;color:var(--ink);font-size:15px;line-height:1.45;margin:2rem 0 20px;display:flex;flex-direction:column;gap:28px}
#genai-census *{box-sizing:border-box}
#genai-census .q-meta{font-size:12.5px;color:var(--ink3);font-variant-numeric:tabular-nums}
#genai-census .q-fig{display:flex;flex-direction:column;gap:10px;border:1px solid var(--rule);border-radius:4px;padding:18px;background:#fff}
#genai-census .q-cap{font-weight:600;font-size:14.5px;color:var(--ink);margin:0}
#genai-census .q-sub{font-size:13px;color:var(--ink2);margin:0}
#genai-census .q-stack{display:flex;gap:2px;height:30px}
#genai-census .q-stack div{border-radius:3px;min-width:3px;display:flex;align-items:center;padding-left:8px;font-size:12.5px;font-weight:600;color:#fff;overflow:hidden;white-space:nowrap}
#genai-census .q-legend{display:flex;flex-wrap:wrap;gap:4px 16px;font-size:12.5px;color:var(--ink2)}
#genai-census .q-legend i{display:inline-block;width:10px;height:10px;border-radius:2px;margin-right:6px}
#genai-census .q-bars{display:flex;flex-direction:column;gap:4px}
#genai-census .q-bar{display:grid;grid-template-columns:minmax(120px,44%) 1fr 28px;gap:10px;align-items:center;width:100%;background:none;border:0;border-radius:3px;padding:4px 6px;font:inherit;font-size:13.5px;color:var(--ink2);text-align:left;cursor:pointer}
#genai-census .q-bar:hover,#genai-census .q-bar:focus-visible{background:var(--navy3);outline:none}
#genai-census .q-bar[aria-pressed="true"]{background:var(--navy3);color:var(--ink);font-weight:600}
#genai-census .q-bar span.t{display:block;height:12px}
#genai-census .q-bar span.t i{display:block;height:100%;background:var(--navy);border-radius:0 3px 3px 0;min-width:2px}
#genai-census .q-bar span.n{text-align:right;font-variant-numeric:tabular-nums;color:var(--ink3);font-size:12.5px}
#genai-census .q-tablewrap{overflow-x:auto;max-height:420px;overflow-y:auto;border:1px solid var(--rule);border-radius:3px}
#genai-census table{border-collapse:collapse;width:100%;min-width:600px;font-size:13px;margin:0}
#genai-census th{position:sticky;top:0;background:var(--navy3);text-align:left;padding:7px 9px;font-size:11.5px;letter-spacing:.04em;text-transform:uppercase;color:var(--ink2);font-weight:600}
#genai-census td{padding:7px 9px;border-top:1px solid var(--rule);vertical-align:top;color:var(--ink2)}
#genai-census td a{color:var(--navy);font-weight:600}
#genai-census .q-pill{display:inline-block;font-size:11px;padding:1px 7px;border-radius:9px;border:1px solid var(--rule);white-space:nowrap;margin:0 3px 3px 0;color:var(--ink2)}
#genai-census .q-pill.no{border-color:var(--orange);color:var(--orange)}
#genai-census .q-row{display:flex;justify-content:space-between;align-items:center;gap:12px;flex-wrap:wrap}
#genai-census .q-clear{font:inherit;font-size:12.5px;border:1px solid var(--rule);background:#fff;border-radius:3px;padding:3px 9px;cursor:pointer;color:var(--navy)}
#genai-census .q-clear[hidden]{display:none}
#genai-census .q-prow{display:grid;grid-template-columns:minmax(130px,34%) 1fr 30px;gap:10px;align-items:center;width:100%;background:none;border:0;border-radius:3px;padding:5px 6px;font:inherit;font-size:14px;color:var(--ink2);text-align:left;cursor:pointer}
#genai-census .q-prow:hover,#genai-census .q-prow:focus-visible{background:var(--navy3);outline:none}
#genai-census .q-prow[aria-pressed="true"]{background:var(--navy3);color:var(--ink);font-weight:600}
#genai-census .q-prow small{display:block;font-size:11px;color:var(--ink3);font-weight:400}
#genai-census .q-prow .seg{display:flex;gap:2px;height:16px}
#genai-census .q-prow .seg i{display:block;height:100%;border-radius:2px;min-width:3px}
#genai-census .q-prow .n{text-align:right;font-variant-numeric:tabular-nums;color:var(--ink3);font-size:12.5px}
#genai-census .q-pill.yes{border-color:var(--navy);color:var(--navy)}
</style>
<div id="genai-census">
  <div class="q-meta" id="gc-meta"></div>
  <div class="q-fig">
    <p class="q-cap">Figure 1. Positions on the discussion paper's eleven proposals</p>
    <p class="q-sub">Unique submitters who take up each proposal, ordered from least to most unqualified support. Bar length is the number of submitters (n at right). Select a row to open that proposal in Figure 2.</p>
    <div class="q-legend" id="gc-legend"></div>
    <div class="q-bars" id="gc-props"></div>
  </div>
</div>
<script>
(function(){
  var D = {"generated": "2026-09-21", "comments_total": 98, "submitters_total": 86, "proposals": [{"key": "two_axis_risk_framework", "name": "Two-axis risk framework", "qs": "Q1-3, Q8", "n": 52, "counts": {"conditional": 47, "oppose": 2, "support": 2, "mixed": 1}}, {"key": "generalist_vs_specialist_distinction", "name": "Generalist vs specialist users", "qs": "Q4", "n": 16, "counts": {"support": 4, "conditional": 10, "oppose": 2}}, {"key": "patient_facing_devices", "name": "Patient-facing devices", "qs": "Q5-6", "n": 39, "counts": {"conditional": 36, "oppose": 1, "support": 1, "mixed": 1}}, {"key": "competency_based_benchmarking", "name": "Competency benchmarking", "qs": "Q7, Q9-10, Q17", "n": 49, "counts": {"conditional": 45, "support": 3, "mixed": 1}}, {"key": "clinical_confirmation_studies", "name": "Clinical confirmation", "qs": "Q11-12, Q14-15", "n": 41, "counts": {"conditional": 35, "support": 6}}, {"key": "synthetic_data", "name": "Synthetic data", "qs": "Q13", "n": 23, "counts": {"conditional": 20, "mixed": 3}}, {"key": "independent_third_party_testing", "name": "Third-party testing", "qs": "Q16", "n": 28, "counts": {"conditional": 22, "mixed": 2, "support": 3, "oppose": 1}}, {"key": "premarket_for_postmarket_tradeoff", "name": "Premarket-for-postmarket trade", "qs": "Q18", "n": 30, "counts": {"conditional": 28, "oppose": 2}}, {"key": "postmarket_monitoring", "name": "Postmarket monitoring", "qs": "Q19-21", "n": 59, "counts": {"conditional": 52, "support": 5, "mixed": 2}}, {"key": "pccp_change_control", "name": "PCCP / change control", "qs": "Q22-24", "n": 49, "counts": {"conditional": 46, "support": 2, "mixed": 1}}, {"key": "foundation_model_maf", "name": "Foundation model MAF", "qs": "Q25", "n": 25, "counts": {"conditional": 22, "support": 2, "oppose": 1}}]};
  var CL = {high_risk_or_high_consequence:'Not for high-risk or high-consequence functions',irreversible_harm:'Harm must be reversible',detectability_or_time_to_detection:'Failure must be detectable before harm accumulates',autonomy_level:'Limited device autonomy',prespecified_triggers_thresholds:'Prespecified triggers and thresholds',monitoring_completeness_or_blind_spots:'Monitoring coverage must be demonstrated',rollback_or_kill_switch:'Rollback or stop mechanism',enforceable_postmarket_commitments:'Monitoring enforceable as a condition of authorization',monitor_independence:'Monitor independent of the device model',vulnerable_or_unable_to_self_report_users:'Not for users who cannot self-report errors',patient_facing:'Stricter for patient-facing devices'};
  var SEG = [['oppose','Oppose','var(--orange)'],['mixed','Mixed','#B9BCC9'],['conditional','Support with changes or conditions','var(--navy2)'],['support','Support','var(--navy)']];
  var PL = {oppose:'Oppose',mixed:'Mixed',conditional:'With changes',support:'Support'};
  var $ = function(id){return document.getElementById(id);};
  var esc = function(s){return String(s).replace(/[&<>"]/g,function(c){return {'&':'&amp;','<':'&lt;','>':'&gt;','"':'&quot;'}[c];});};
  $('gc-meta').textContent = 'Data as of '+D.generated+' · docket FDA-2026-N-7874 · '+D.comments_total+' comments from '+D.submitters_total+' submitters';
  SEG.forEach(function(s){var l=document.createElement('span');l.innerHTML='<i style="background:'+s[2]+'"></i>'+s[1];$('gc-legend').appendChild(l);});
  var props = D.proposals.slice().sort(function(a,b){return (a.counts.support||0)/a.n-(b.counts.support||0)/b.n || (b.counts.oppose||0)-(a.counts.oppose||0);});
  var max = Math.max.apply(null,props.map(function(p){return p.n;}));
  props.forEach(function(p){
    var bt=document.createElement('button');bt.type='button';bt.className='q-prow';bt.id='gc-p-'+p.key;bt.dataset.k=p.key;
    var seg=SEG.map(function(s){var n=p.counts[s[0]]||0;return n?'<i title="'+s[1]+': '+n+'" style="background:'+s[2]+';width:'+(100*n/max)+'%"></i>':'';}).join('');
    bt.innerHTML='<span>'+esc(p.name)+'<small>'+p.qs+'</small></span><span class="seg">'+seg+'</span><span class="n">'+p.n+'</span>';
    bt.setAttribute('aria-label',p.name+': '+SEG.map(function(s){return s[1]+' '+(p.counts[s[0]]||0);}).join(', '));
    bt.addEventListener('click',function(){var s=document.getElementById('g18-sel'); if(s){s.value=p.key;s.dispatchEvent(new Event('change'));s.scrollIntoView({behavior:'smooth',block:'center'});}});$('gc-props').appendChild(bt);});
})();
</script>

<!-- GenAI docket census, Figure 2 (requested changes by proposal). Generated by build_embed.py. Replace this Notion HTML embed with the regenerated file. -->
<style>
#genai-q18{--navy:#2D3F86;--navy2:#8F9BD0;--navy3:#E6E9F6;--orange:#E85036;--ink:#161B33;--ink2:#4A5172;--ink3:#7B819D;--rule:#D9DCEA;color:var(--ink);font-size:15px;line-height:1.45;margin:0 0 2rem;display:flex;flex-direction:column;gap:28px}
#genai-q18 *{box-sizing:border-box}
#genai-q18 .q-meta{font-size:12.5px;color:var(--ink3);font-variant-numeric:tabular-nums}
#genai-q18 .q-fig{display:flex;flex-direction:column;gap:10px;border:1px solid var(--rule);border-radius:4px;padding:18px;background:#fff}
#genai-q18 .q-cap{font-weight:600;font-size:14.5px;color:var(--ink);margin:0}
#genai-q18 .q-sub{font-size:13px;color:var(--ink2);margin:0}
#genai-q18 .q-stack{display:flex;gap:2px;height:30px}
#genai-q18 .q-stack div{border-radius:3px;min-width:3px;display:flex;align-items:center;padding-left:8px;font-size:12.5px;font-weight:600;color:#fff;overflow:hidden;white-space:nowrap}
#genai-q18 .q-legend{display:flex;flex-wrap:wrap;gap:4px 16px;font-size:12.5px;color:var(--ink2)}
#genai-q18 .q-legend i{display:inline-block;width:10px;height:10px;border-radius:2px;margin-right:6px}
#genai-q18 .q-bars{display:flex;flex-direction:column;gap:4px}
#genai-q18 .q-bar{display:grid;grid-template-columns:minmax(120px,44%) 1fr 28px;gap:10px;align-items:center;width:100%;background:none;border:0;border-radius:3px;padding:4px 6px;font:inherit;font-size:13.5px;color:var(--ink2);text-align:left;cursor:pointer}
#genai-q18 .q-bar:hover,#genai-q18 .q-bar:focus-visible{background:var(--navy3);outline:none}
#genai-q18 .q-bar[aria-pressed="true"]{background:var(--navy3);color:var(--ink);font-weight:600}
#genai-q18 .q-bar span.t{display:block;height:12px}
#genai-q18 .q-bar span.t i{display:block;height:100%;background:var(--navy);border-radius:0 3px 3px 0;min-width:2px}
#genai-q18 .q-bar span.n{text-align:right;font-variant-numeric:tabular-nums;color:var(--ink3);font-size:12.5px}
#genai-q18 .q-tablewrap{overflow-x:auto;max-height:420px;overflow-y:auto;border:1px solid var(--rule);border-radius:3px}
#genai-q18 table{border-collapse:collapse;width:100%;min-width:600px;font-size:13px;margin:0}
#genai-q18 th{position:sticky;top:0;background:var(--navy3);text-align:left;padding:7px 9px;font-size:11.5px;letter-spacing:.04em;text-transform:uppercase;color:var(--ink2);font-weight:600}
#genai-q18 td{padding:7px 9px;border-top:1px solid var(--rule);vertical-align:top;color:var(--ink2)}
#genai-q18 td a{color:var(--navy);font-weight:600}
#genai-q18 .q-pill{display:inline-block;font-size:11px;padding:1px 7px;border-radius:9px;border:1px solid var(--rule);white-space:nowrap;margin:0 3px 3px 0;color:var(--ink2)}
#genai-q18 .q-pill.no{border-color:var(--orange);color:var(--orange)}
#genai-q18 .q-row{display:flex;justify-content:space-between;align-items:center;gap:12px;flex-wrap:wrap}
#genai-q18 .q-clear{font:inherit;font-size:12.5px;border:1px solid var(--rule);background:#fff;border-radius:3px;padding:3px 9px;cursor:pointer;color:var(--navy)}
#genai-q18 .q-clear[hidden]{display:none}
#genai-q18 .q-prow{display:grid;grid-template-columns:minmax(130px,34%) 1fr 30px;gap:10px;align-items:center;width:100%;background:none;border:0;border-radius:3px;padding:5px 6px;font:inherit;font-size:14px;color:var(--ink2);text-align:left;cursor:pointer}
#genai-q18 .q-prow:hover,#genai-q18 .q-prow:focus-visible{background:var(--navy3);outline:none}
#genai-q18 .q-prow[aria-pressed="true"]{background:var(--navy3);color:var(--ink);font-weight:600}
#genai-q18 .q-prow small{display:block;font-size:11px;color:var(--ink3);font-weight:400}
#genai-q18 .q-prow .seg{display:flex;gap:2px;height:16px}
#genai-q18 .q-prow .seg i{display:block;height:100%;border-radius:2px;min-width:3px}
#genai-q18 .q-prow .n{text-align:right;font-variant-numeric:tabular-nums;color:var(--ink3);font-size:12.5px}
#genai-q18 .q-pill.yes{border-color:var(--navy);color:var(--navy)}
#genai-q18 .q-grp{font-size:12px;letter-spacing:.05em;text-transform:uppercase;color:var(--ink3);font-weight:600;margin:10px 0 2px 6px}
#genai-q18 select{font:inherit;font-size:14px;padding:6px 10px;border:1px solid var(--rule);border-radius:3px;background:#fff;color:var(--ink);max-width:100%}
#genai-q18 select:focus-visible{outline:2px solid var(--navy)}
#genai-q18 label{font-size:13px;color:var(--ink2);margin-right:8px}
</style>
<div id="genai-q18">
  <div class="q-fig">
    <p class="q-cap">Figure 2. What submitters ask FDA to change, by proposal</p>
    <div><label for="g18-sel">Proposal</label><select id="g18-sel"></select></div>
    <p class="q-sub" id="g18-sub"></p>
    <div class="q-bars" id="g18-bars"></div>
    <div class="q-row"><p class="q-cap" id="g18-tcap" style="margin-top:8px"></p><button type="button" class="q-clear" id="g18-clear" hidden>Show all</button></div>
    <div class="q-tablewrap" style="max-height:320px"><table><thead><tr><th>Submitter</th><th>Position</th><th>Summary of position</th></tr></thead><tbody id="g18-body"></tbody></table></div>
    <p class="q-sub" id="g18-meta"></p>
  </div>
</div>
<script>
(function(){
  var D = {"generated": "2026-09-21", "comments_total": 98, "submitters_total": 86, "proposals": [{"key": "two_axis_risk_framework", "name": "Two-axis risk framework", "qs": "Q1-3, Q8", "n": 52, "counts": {"conditional": 47, "oppose": 2, "support": 2, "mixed": 1}, "rows": [{"id": "0003", "p": "conditional", "g": "Right organizing principle, but needs added dimensions (third-party model dependency, reversibility, temporal urgency), a numeric-parameter safe-harbor presumption, and a pathway decision matrix."}, {"id": "0005", "p": "conditional", "g": "Two-axis matrix is useful as a communication layer but insufficient as the full risk model; add a broader boundary ledger and hard gates."}, {"id": "0006", "p": "conditional", "g": "Framework is workable but should add counterfactual care environment and duration of reliance as dimensions; judge directiveness by substance and publish worked examples."}, {"id": "0009", "p": "conditional", "g": "Calls the framework a reasonable heuristic and asks for three added dimensions: reversibility and time to consequence, depth of downstream human review, and traceability to source evidence."}, {"id": "0010", "p": "conditional", "g": "Axes are necessary but omit detectability of incorrect outputs; should be reconciled with ISO 14971, treat directiveness as a monitored property, and be paired with a published evidence mapping."}, {"id": "0012", "p": "conditional", "g": "Supports the two-axis framework provided scribing, EHR screening and supervised imaging second-reads sit in lower-risk categories, with heavy scrutiny reserved for unmonitored autonomous execution."}, {"id": "0013", "p": "conditional", "g": "The two axes are necessary but not sufficient; they capture severity but omit probability of undetected error. Add detectability and reconcile with ISO 14971."}, {"id": "0015", "p": "conditional", "g": "Framework is a sound start but insufficient; expand to four axes adding degree of human oversight at output delivery and traceability of output to source data."}, {"id": "0017", "p": "conditional", "g": "Two axes are necessary but insufficient; add irreversibility, time to harm, detection/rescue opportunity, system authority, human review quality and environmental dependence."}, {"id": "0019", "p": "oppose", "g": "Says the two-axis framework does not capture GenAI risk and is not needed; stresses privacy, consent, disclaimers and continued reliance on human mental health professionals."}, {"id": "0020", "p": "conditional", "g": "Clinically intuitive starting point but has too few dimensions; add error detectability, exposure frequency, irreversibility, and tie tiers to evidence requirements."}, {"id": "0021", "p": "conditional", "g": "Consequence axis wrongly limited to incorrect outputs; must add omission, harmful affirmation, relational features as risk modifiers, and pediatric use as multiplier."}, {"id": "0024", "p": "conditional", "g": "Accepts the two-axis approach but wants additional dimensions considered: model autonomy, user ability to catch errors, reversibility, population vulnerability, and clinical and information context."}, {"id": "0025", "p": "conditional", "g": "Useful starting point, but risk must also cover omission, latent clinical need, inappropriate reassurance, foreseeable influence on patient behavior, time sensitivity, reversibility and cumulative multi-turn effects."}, {"id": "0027", "p": "conditional", "g": "For agentic devices, consequence assessment should capture actions taken without an opportunity for human reliance to be withheld, including reversibility and authorization controls."}, {"id": "0028", "p": "conditional", "g": "Supports the framework as a starting heuristic, not a scoring tool, and wants reversibility, downstream safeguards, time pressure and traceability added as explicit modifiers."}, {"id": "0029", "p": "support", "g": "Framework appropriately considers device activity independence and consequence of incorrect output; required evidence for continued reliance should increase with both."}, {"id": "0030", "p": "conditional", "g": "Framework useful for initial characterization but should be a screening tool, not a classification matrix; device-level risk analysis with error-progression factors must follow."}, {"id": "0032", "p": "conditional", "g": "Accepts the activity-and-consequence framework but wants connectivity dependence, time-to-harm during unavailability, fallback capability, detectability of degradation and common-mode exposure added as explicit risk modifiers."}, {"id": "0035", "p": "conditional", "g": "Two axes are appropriate for inherent risk but should be modified by a structured profile of reversibility, time to harm, oversight, traceability, state persistence, scale, and uncertainty visibility."}, {"id": "0037", "p": "conditional", "g": "Consequence axis is appropriate, but the activity axis is insufficient without architectural connectedness and agentic coupling; directiveness should modify probability of harm using existing benefit-risk vocabulary."}, {"id": "0038", "p": "conditional", "g": "Two-axis model is a good starting point but omits upstream and lifecycle risks; proposes three-layer model retaining the two axes as Layer 1."}, {"id": "0039", "p": "conditional", "g": "Framework should be expanded to add autonomy, detectability, reversibility, human oversight, patient vulnerability and severity of harm as dimensions."}, {"id": "0042", "p": "conditional", "g": "Calls the supervised/autonomous distinction correct but unverifiable from current records; asks that attribution traceability be added as a risk dimension raising risk for devices that cannot preserve it."}, {"id": "0043", "p": "conditional", "g": "The framework can work, but probability and consequence are hard to measure, traceability is infeasible for opaque models, and human-in-the-loop status and detectability should be added."}, {"id": "0045", "p": "conditional", "g": "Says two axes are insufficient and recommends five: consequence, clinical agency, reversibility, time-to-harm and human recoverability, with evidence scaling to agency."}, {"id": "0046", "p": "conditional", "g": "Endorses the framework but wants a named human-in-the-loop and reversibility treated as explicit risk-lowering variables, and the 520(o)(1)(E) independent-review boundary named as a lower-risk class."}, {"id": "0047", "p": "conditional", "g": "Welcomes risk-based regulation scaled to decision seriousness, patient vulnerability and error consequences, but wants licensed-clinician review mandatory for diagnosis, prescribing, triage and escalation decisions."}, {"id": "0048", "p": "conditional", "g": "Finds the activity and consequence axes useful but wants modifiers such as reversibility, time to correction, detectability, traceability and safeguard independence; directiveness judged by observed behavior."}, {"id": "0049", "p": "conditional", "g": "Retain the two axes as the primary heuristic, but treat traceability, reversibility, time pressure, safeguards, user capability and execution authority as risk modifiers; scale evidence continuously."}, {"id": "0051", "p": "conditional", "g": "Endorses the two axes; wants extra dimensions handled as documented modifiers on the consequences axis, adding user detectability of errors, plus a published evidence gradient and directiveness rubric."}, {"id": "0052", "p": "conditional", "g": "Useful starting point, but FDA should separate initial regulatory criticality stratification from ISO 14971 product risk, add modifiers, and credit only defined, enforceable safeguards, stratifying conservatively."}, {"id": "0053", "p": "conditional", "g": "Reversibility and traceability should be represented, with a physical-output modifier presumptively placing implant-geometry-generating functions at the high-consequence end."}, {"id": "0054", "p": "conditional", "g": "Supports the two-axis concept as a foundation but asks for four explicit modifiers: evidence traceability, independent reviewability, reversibility and time to harm."}, {"id": "0055", "p": "oppose", "g": "Says the two-axis framework is insufficient; CDRH should keep probability-and-severity hazard analysis, divide output into general versus patient-specific, and not let the axes set evidence."}, {"id": "0056", "p": "conditional", "g": "Wants the risk assessment to add whether errors are independently detectable by the intended user, reversibility before harm, and realistic availability of specialist escalation."}, {"id": "0057", "p": "conditional", "g": "Calls the two-axis framework sound but wants human oversight treated as an empirical property demonstrated in representative use, with automation bias modifying activity-axis position."}, {"id": "0058", "p": "conditional", "g": "Calls the two-axis framework sound but wants human oversight treated as an empirical property demonstrated in representative use, with automation bias modifying activity-axis position."}, {"id": "0065", "p": "conditional", "g": "The two axes applied turn by turn do not register migration to directive output; directiveness is a variable driven by user affect, not a modifier of intended use."}, {"id": "0067", "p": "conditional", "g": "The two axes restate model influence and decision consequence from existing FDA credibility frameworks; FDA should adopt those terms or declare equivalence, and let the tier drive evidence."}, {"id": "0076", "p": "conditional", "g": "Supports the two-axis framework as a sound heuristic but asks that data exposure and privacy be added as a third axis or mandatory consequence modifier."}, {"id": "0080", "p": "conditional", "g": "Supports the framework and asks FDA to use it to declare informational, non-directive functions non-devices or under enforcement discretion, and to reconcile it with the wellness policy."}, {"id": "0082", "p": "conditional", "g": "Finds the framework apt but wants time pressure and downstream safeguards made first-class dimensions, protocol fidelity as the primary risk modifier, and a protocol-delivery sub-category."}, {"id": "0086", "p": "conditional", "g": "Wants the risk framework to assess function plus real conditions of use, including reversibility, urgency, vulnerability, error detection, traceability, scale, staffing and workflow."}, {"id": "0087", "p": "conditional", "g": "Calls the two-axis framework a useful starting point but wants opportunity for meaningful human review and traceability to source evidence made explicit risk factors."}, {"id": "0088", "p": "mixed", "g": "Says source traceability belongs in the risk picture and that hallucination control cannot be shown by regex checks over a fixed entity list."}, {"id": "0089", "p": "conditional", "g": "The risk framework should distinguish agentic devices by reversibility and consequence of autonomous actions, with required evidence of human authorization scaling accordingly."}, {"id": "0091", "p": "support", "g": "Calls the two-axis framing a reasonable basis for choosing among confirmation approaches of increasing rigour and would not add a third axis."}, {"id": "0093", "p": "conditional", "g": "Supports the two axes as an organizing heuristic; asks FDA to add change dynamics as a formal dimension and treat source attribution as risk-lowering."}, {"id": "0094", "p": "conditional", "g": "Supports the two-axis framework as a heuristic but wants evidentiary traceability added as a dimension, deployment-context modifiers, and harmonization with HHS high-impact AI classification."}, {"id": "0097", "p": "conditional", "g": "Wants risk assessment to weigh error severity, a clear regulatory distinction between augmenting and replacing physician judgment, and safeguards that increase with autonomy."}, {"id": "0098", "p": "conditional", "g": "The activity axis places all HCP-supervised action at one point regardless of oversight quality; asks for an oversight-quality modifier based on four expert-in-the-loop criteria."}, {"id": "0099", "p": "conditional", "g": "The two axes identify the right variables, but reversibility should be an explicit gating factor evaluated before fixing a function's final risk position."}, {"id": "0100", "p": "conditional", "g": "Supports the two-axis framework and asks FDA to add traceability and reversibility as formal risk modifiers and to clarify the directiveness continuum with worked examples."}], "themes": [{"name": "Add reversibility and time-to-harm as risk modifiers", "ids": ["0003", "0005", "0009", "0017", "0020", "0024", "0025", "0027", "0028", "0030", "0035", "0039", "0045", "0046", "0048", "0049", "0053", "0054", "0056", "0086", "0089", "0099", "0100"], "n": 23}, {"name": "Add output-to-source traceability as a risk dimension", "ids": ["0005", "0009", "0015", "0028", "0035", "0042", "0048", "0049", "0053", "0054", "0087", "0088", "0093", "0094", "0100"], "n": 15}, {"name": "Add user detectability of errors as a risk dimension", "ids": ["0010", "0013", "0020", "0024", "0030", "0039", "0043", "0048", "0051", "0056", "0086"], "n": 10}, {"name": "Judge directiveness by observed behavior, not labels or disclaimers", "ids": ["0006", "0038", "0045", "0048", "0049", "0057", "0058", "0076"], "n": 7}, {"name": "Add patient vulnerability, including pediatric use, as a risk modifier", "ids": ["0020", "0021", "0024", "0039", "0047", "0086"], "n": 6}, {"name": "Credit human oversight only when demonstrably effective", "ids": ["0017", "0057", "0058", "0087", "0098", "0099"], "n": 5}, {"name": "Publish a table mapping risk tiers to required evidence", "ids": ["0010", "0020", "0051", "0093"], "n": 4}]}, {"key": "generalist_vs_specialist_distinction", "name": "Generalist vs specialist users", "qs": "Q4", "n": 16, "counts": {"support": 4, "conditional": 10, "oppose": 2}, "rows": [{"id": "0003", "p": "support", "g": "Endorses distinguishing generalist and specialist users through competency-specific labeling, intended use, and validation in the intended user population."}, {"id": "0005", "p": "conditional", "g": "Risk should rest on task-to-competency mismatch, not professional title; validate generalist review capability and device specialist-escalation behavior."}, {"id": "0009", "p": "conditional", "g": "Says the patient, generalist and specialist division omits trained non-clinical health workers, and asks for a third user tier that accounts for supervisory structure."}, {"id": "0010", "p": "conditional", "g": "Distinction is real but unenforceable through labeling; risk should be assessed against foreseeable use with credit only for technical user-scope enforcement."}, {"id": "0013", "p": "conditional", "g": "The distinction is real but unenforceable through labeling; risk should be assessed against foreseeable use, with credit only for demonstrated technical enforcement of user scope."}, {"id": "0025", "p": "conditional", "g": "AI can bring specialist knowledge to generalists, but must be evidence-grounded, recognize limits of generalist practice, flag when specialist involvement is needed, and convey uncertainty."}, {"id": "0028", "p": "conditional", "g": "Risk should reflect the verifiable gap between the function's knowledge domain and intended users, not a blanket assumption that generalist use is riskier."}, {"id": "0037", "p": "oppose", "g": "Rejects generalist-versus-specialist and other credential categories as coarse proxies; would replace them with demonstrated, device-specific competency affirmation scaled to risk."}, {"id": "0043", "p": "support", "g": "Supports distinguishing specialist-directed outputs through explicit modifiers warning generalists about the intended specialist audience."}, {"id": "0045", "p": "oppose", "g": "Rejects a generalist-versus-specialist label; the comparator should be defined by the clinical task in its intended environment and set prospectively."}, {"id": "0049", "p": "conditional", "g": "Treat the distinction as a context modifier, not a categorical label; increase safeguards where safe use needs specialist interpretation the user lacks, and build confirmation into workflow."}, {"id": "0051", "p": "conditional", "g": "Accepts the distinction if handled through labeling of expected user competency, testing with the least specialized intended user, and tested referral-prompting behavior."}, {"id": "0052", "p": "conditional", "g": "The distinction may be relevant, but any modifying effect should rest on assured competence, information and workflow conditions, not professional designation; generalists may see the broader patient picture."}, {"id": "0055", "p": "support", "g": "Agrees risk rises when users lack needed specialist knowledge; the intended user is part of intended use, and specialist and generalist devices need separate confirmation."}, {"id": "0056", "p": "support", "g": "Specialty knowledge should inform risk when detecting errors requires it; specialty functions need qualified specialist evaluation even when generalists use them."}, {"id": "0093", "p": "conditional", "g": "Recommends separating whether safe use depends on specialist contextualization from whether deployment assumes specialist availability, tested through elements S.3 and E.1."}, {"id": "0098", "p": "conditional", "g": "Agrees with the Question 4 concern and says the same domain-match gap applies to human reviewers in supervision or adjudication roles."}], "themes": [{"name": "Base risk on task-competency gap, not professional title", "ids": ["0005", "0028", "0037", "0052"], "n": 4}, {"name": "Test the device's specialist-referral and deferral behavior", "ids": ["0005", "0025", "0051", "0093"], "n": 4}, {"name": "Test error detection with the least specialized intended users", "ids": ["0005", "0051"], "n": 2}]}, {"key": "patient_facing_devices", "name": "Patient-facing devices", "qs": "Q5-6", "n": 39, "counts": {"conditional": 36, "oppose": 1, "support": 1, "mixed": 1}, "rows": [{"id": "0003", "p": "conditional", "g": "Supports differential treatment of patient-facing devices but warns against overreach on general informational tools; assess multi-turn risk at interaction level; weight under-escalation more heavily."}, {"id": "0005", "p": "conditional", "g": "Do not infer safety from patient/HCP labels; measure user comprehension and error detection, evaluate full conversational trajectories, and keep under- and over-escalation separate."}, {"id": "0006", "p": "conditional", "g": "Patient-facing should not be a per se risk elevation; credit verifiable behaviors, assess conversations at trajectory level via a conversational envelope, and prespecify asymmetric escalation tolerances."}, {"id": "0009", "p": "conditional", "g": "Agrees both escalation error directions matter and asks FDA to recognize routing to an accountable clinical channel as a distinct category, with both error rates reported separately."}, {"id": "0010", "p": "conditional", "g": "Do not categorically elevate patient-facing risk; focus on verification path. For multi-turn devices, make the session the unit of analysis; prespecify asymmetric escalation criteria."}, {"id": "0013", "p": "conditional", "g": "Opposes categorical risk elevation for patient-facing functions; wants session-level intended use, trajectory-based adversarial evaluation, and prespecified asymmetric escalation criteria."}, {"id": "0017", "p": "conditional", "g": "Multi-turn and agentic systems should be evaluated across realistic trajectories, not isolated outputs; under- and over-escalation should be evaluated separately."}, {"id": "0019", "p": "oppose", "g": "Rejects patient-facing AI in mental health; patients should keep using human professionals, and no escalation trade-offs are acceptable in clinical context."}, {"id": "0020", "p": "support", "g": "Praises treating multi-turn conversations as a unit of risk and weighing both under- and over-escalation errors as clinically mature."}, {"id": "0021", "p": "conditional", "g": "Supports patient-facing devices in principle but multi-turn must span days/months and sessions; escalation must score a third error, 'escalating badly', with proven handoff."}, {"id": "0023", "p": "conditional", "g": "Patient-facing AI information and AI therapists pose severe risks; should require human oversight and FDA regulation of AI therapists as medical devices."}, {"id": "0025", "p": "conditional", "g": "Patient-facing AI carries distinct risks; FDA should assess whole multi-turn encounters, weigh both escalation errors, and favor triage grounded in curated evidence-based guidelines rather than internet-derived standards."}, {"id": "0026", "p": "conditional", "g": "Care escalation functions should prioritize sensitivity and minimize under-escalation; over-escalation acceptable if an intermediate pathway exists; manufacturers justify the balance."}, {"id": "0027", "p": "conditional", "g": "Conversational devices require direct trajectory-level testing with distinct acceptance criteria because single-turn evaluations cannot establish resistance to multi-turn escalation."}, {"id": "0028", "p": "conditional", "g": "Supports differentiated treatment but opposes automatic elevation of patient-facing functions; backs trajectory-based assessment and sponsor-justified escalation operating points."}, {"id": "0030", "p": "conditional", "g": "Agrees user evaluation ability affects risk, but patient-facing use should not presume higher risk; adaptive comprehension assessment can serve as an optional risk control."}, {"id": "0035", "p": "conditional", "g": "Patient-facing status is a context modifier, not presumptively unacceptable; directiveness and escalation risk should be evaluated across conversational trajectories with separate under- and over-escalation thresholds."}, {"id": "0037", "p": "conditional", "g": "Objects to presuming patients lack domain knowledge; wants competency affirmation scaled to risk, high oversight for devices that can migrate to directive output, and escalation-error weights tied to user competency."}, {"id": "0038", "p": "conditional", "g": "Patient-facing outputs should be plain-language, non-directive, with optional full data access and safeguards; disclaimers do not reduce directiveness."}, {"id": "0043", "p": "conditional", "g": "Patient delivery is unequivocally higher risk and patient empowerment is a poor goal; conversational devices should detect emergent advice and withhold it from patients with a warning."}, {"id": "0045", "p": "conditional", "g": "Patient-facing AI should not be penalized by default but needs distinct safeguards; wants trajectory testing and sponsor-defined context-specific escalation error costs."}, {"id": "0047", "p": "conditional", "g": "Allows AI to inform clinical decisions, including triage and escalation, provided licensed clinicians review and approve them and vulnerable patients receive stronger protections."}, {"id": "0048", "p": "conditional", "g": "Calls for multi-turn evaluation covering conversational context, escalation, refusal, interruption, and safe stopping; does not discuss patient-specific risks."}, {"id": "0049", "p": "conditional", "g": "Patient-facing status should not automatically mean higher risk; safeguards depend on context. Evaluate multi-turn trajectories with transition testing and report under- and over-escalation separately."}, {"id": "0051", "p": "conditional", "g": "Supports higher risk placement for patient-facing functions only conditionally, offset by demonstrated mitigations; proposes behavioral envelopes for multi-turn devices and decision-analytic weighting of escalation errors."}, {"id": "0052", "p": "conditional", "g": "Patient-facing distinction matters, but criticality should turn on enforceable safeguards in the use environment; multi-turn devices take the highest foreseeable criticality; escalation should be graded, not binary."}, {"id": "0054", "p": "conditional", "g": "Patient-facing systems should not be high risk by default; risk should follow functional category, and under- and over-escalation should be reported separately with severity weighting."}, {"id": "0055", "p": "conditional", "g": "Patient-facing devices can carry more risk but patients should not be presumed incapable; asks for scope limits, escalation, whole-conversation testing, deterministic controls and separate limits for severe under-escalation."}, {"id": "0061", "p": "conditional", "g": "Does not exclude clinical deployment, but demands safeguards accounting for users unable to contest outputs and inspect hidden assertions about themselves."}, {"id": "0065", "p": "conditional", "g": "Multi-turn devices need not be excluded, but risk class should follow the most directive output reachable on a trajectory, tested by varying user affect with clinical content constant."}, {"id": "0067", "p": "conditional", "g": "On care escalation, the two error directions should not be collapsed into one accuracy number; sponsors should prespecify the weighting and show performance across the trade-off."}, {"id": "0076", "p": "conditional", "g": "Supports patient-facing access with transparency and recourse rather than restriction; escalation trade-offs should be prespecified and weighted toward avoiding under-escalation."}, {"id": "0080", "p": "conditional", "g": "Consumer-facing functions that contextualize a user's own wearable data and suggest consulting a clinician are low risk and should be enabled without active device regulation."}, {"id": "0082", "p": "conditional", "g": "Wants lay-user design safeguards counted as mitigants and escalation in protocol-delivery functions scored as scenario-specific ordering rather than a trade-off between error rates."}, {"id": "0084", "p": "conditional", "g": "On care escalation, both error directions carry real costs; thresholds should be justified against the specific population and care access context, not one error rate."}, {"id": "0088", "p": "mixed", "g": "Says multi-turn risk includes failing to use already-collected context, and that missing a key history question is an error analogous to under-escalation."}, {"id": "0091", "p": "conditional", "g": "Multi-turn evaluation needs a prespecified, validated observation window with timeliness separated from recovery; under- and over-escalation need separate outcomes and denominators."}, {"id": "0093", "p": "conditional", "g": "Supports placing patient-facing triage, dose or escalation functions higher on the consequences axis, with mitigations, a declared conversation-state envelope, and paired escalation-error endpoints."}, {"id": "0094", "p": "conditional", "g": "Favors allowing patient-facing information without blanket restriction, mitigated by a worker-extender certification model that anchors accountability to a licensed human role."}, {"id": "0098", "p": "conditional", "g": "Confirms multi-turn drift as an observed risk and asks for cumulative-state drift testing and fail-closed escalation testing under fault conditions."}, {"id": "0100", "p": "conditional", "g": "Supports patient-facing functions and bidirectional escalation assessment; opposes a categorical higher-risk presumption, asking that risk turn on verifiable safeguards and prespecified escalation operating points."}], "themes": [{"name": "Evaluate whole multi-turn conversation trajectories, not single turns", "ids": ["0003", "0005", "0006", "0010", "0013", "0017", "0025", "0027", "0028", "0035", "0045", "0048", "0049", "0055", "0098"], "n": 14}, {"name": "Report under-escalation and over-escalation rates separately", "ids": ["0005", "0009", "0017", "0049", "0051", "0054", "0055", "0091", "0093", "0100"], "n": 10}, {"name": "Do not presume patient-facing functions are higher risk", "ids": ["0006", "0010", "0013", "0028", "0030", "0045", "0049", "0054", "0100"], "n": 8}, {"name": "Define intended use as a bounded conversational envelope", "ids": ["0006", "0028", "0035", "0051", "0065", "0093"], "n": 6}, {"name": "Require prespecified, clinically justified escalation thresholds per deployment context", "ids": ["0010", "0013", "0035", "0049", "0067", "0084", "0100"], "n": 6}, {"name": "Weight under-escalation errors more heavily than over-escalation", "ids": ["0003", "0026", "0076"], "n": 3}, {"name": "Require human oversight or a defined route to a human", "ids": ["0023", "0047", "0076"], "n": 3}]}, {"key": "competency_based_benchmarking", "name": "Competency benchmarking", "qs": "Q7, Q9-10, Q17", "n": 49, "counts": {"conditional": 45, "support": 3, "mixed": 1}, "rows": [{"id": "0003", "p": "conditional", "g": "Strongly supports competency-based approach, but requires specialty domain modules, a sequestered NIST-stewarded benchmark repository, third-party validation of sponsor benchmarks, and multimodal companion guidance."}, {"id": "0005", "p": "conditional", "g": "Competency approach useful if benchmark, clinical confirmation and monitoring form one closure chain tied to deployed configuration; add construct-validity and provenance gates."}, {"id": "0006", "p": "support", "g": "The competency-based model is sound and maps naturally onto how medicine already evaluates human clinical judgment."}, {"id": "0009", "p": "conditional", "g": "Requests that communication and comprehension benchmarking specifically evaluate trained non-clinical health workers."}, {"id": "0010", "p": "conditional", "g": "Supports competency framing as structure but not its evidentiary permissiveness; add adversarial, citation-fidelity and edge-of-scope elements; treat public benchmarks as screens; lock and escrow sponsor benchmarks."}, {"id": "0012", "p": "support", "g": "Strongly supports the competency-based model and asks that Safety Escalation and Quantitative Analysis competencies be prioritized for blood lab analysis."}, {"id": "0013", "p": "conditional", "g": "Supports the competency framing as organizing structure but not the licensure analogy's evidentiary permissiveness; wants benchmark gaps filled and public benchmarks treated only as screens."}, {"id": "0017", "p": "conditional", "g": "Supports competency-based model if competency refers to the complete deployed device in its environment; benchmarks need construct validity, protected test sets, robustness and rare-event testing."}, {"id": "0018", "p": "conditional", "g": "Endorses the competency-based framework but says it should add cybersecurity, standalone human factors validation, and layered RAG validation expectations."}, {"id": "0019", "p": "conditional", "g": "Likes the competency-based approach and benchmarking structure but says it needs fleshing out, a pilot study, clinical trials, informed consent, privacy protections and public input first."}, {"id": "0020", "p": "conditional", "g": "Promising and broader than accuracy testing, but risks becoming a surrogate endpoint; for high-risk uses competency must not substitute for clinical utility evidence."}, {"id": "0021", "p": "conditional", "g": "Benchmarking is similar to her lab's approach but must test reasonably foreseeable use and score high-consequence failures separately with minimum floors, no pooled averaging."}, {"id": "0024", "p": "conditional", "g": "Finds the competency-based approach valuable if FDA translates capability into task-specific evidentiary descriptions and benchmarks cover hard, out-of-distribution, adversarial and ambiguous cases plus subgroups."}, {"id": "0025", "p": "conditional", "g": "Appropriate if competency is defined broadly to include knowledge, reasoning, recognition and restraint, plus a new competency for detecting latent clinical need in administrative interactions, tested with realistic messy patients."}, {"id": "0026", "p": "conditional", "g": "Supports benchmarking safety behavior but wants a standardized out-of-domain dataset and predefined deferral threshold, also assessing over-deferral."}, {"id": "0027", "p": "conditional", "g": "Calls the competency structure sound and proposes no new elements, but wants constituent failure-mode resolution, subgroup testing on fine-tuned artifacts, and benchmark exposure analysis for Safety elements."}, {"id": "0028", "p": "conditional", "g": "Supports benchmarking followed by clinical confirmation, with an added interoperability element, embedded subgroup safety criteria, and construct-validity and independence requirements for benchmarks."}, {"id": "0032", "p": "conditional", "g": "Says the benchmarking categories omit cybersecurity, external-service dependence, maintainability and lifecycle continuity; proposes three new elements (S.4, R.3, R.4) and dependency-failure testing."}, {"id": "0035", "p": "conditional", "g": "Supports competency-based evaluation if results bind to an identified system configuration, use distributional testing, pair with deterministic safety invariants, and add an assurance-integrity element; benchmark validity needs an evidence package."}, {"id": "0042", "p": "conditional", "g": "Endorses benchmarking human-oversight checkpoints and proposes additional agentic-device acceptance criteria testing emitted attribution, explicit absence of human authorization, and attribution strength."}, {"id": "0043", "p": "conditional", "g": "Endorses the approach as covering most submissions, but finds the single agentic element A.1 inadequate and wants FDA to give sponsors a starting data framework."}, {"id": "0045", "p": "conditional", "g": "Calls benchmarking plus confirmation the right foundation but wants it extended into a nine-stage licensure pipeline, four added competency domains and sequestered regulatory benchmarks."}, {"id": "0046", "p": "conditional", "g": "Calls the clinician-credentialing analogy apt and the escalation and deferral competencies right, but wants the model benchmarked alone and the human role framed as accountability, not accuracy."}, {"id": "0048", "p": "conditional", "g": "Wants competency shown by structured, claim-specific, traceable evidence rather than one benchmark score, and says public benchmark performance should not be presumed to predict clinical performance."}, {"id": "0049", "p": "conditional", "g": "The competency-based approach is useful and appropriate but necessary, not sufficient; it should be complemented by system-assurance evidence, benchmark records, construct-validity justification and architecture-neutral evaluation."}, {"id": "0051", "p": "conditional", "g": "Calls the approach sound but asks for added elements on temporal validity, input integrity and non-English performance, and predictive-validity evidence with independent sequestered test sets."}, {"id": "0052", "p": "conditional", "g": "Supports the competency-based approach but cautions against equating clinician and GenAI competency; wants dynamic, risk-based evaluation suites, performance floors for critical failures, and staged deployment with evidence gates."}, {"id": "0053", "p": "conditional", "g": "Supports evaluating the final configured device, but for physical-output functions that unit should be the validated production system, benchmarked against a prospectively defined generative design envelope."}, {"id": "0054", "p": "conditional", "g": "Supports the competency-based approach provided competency attaches to the complete deployed system, and asks that evidence fidelity and provenance be added as an explicit competency."}, {"id": "0055", "p": "conditional", "g": "Calls the competency approach useful and necessary, but benchmarking should only gate clinical confirmation, test the final device configuration, and carry separate limits for high-consequence failures."}, {"id": "0056", "p": "conditional", "g": "Calls the structure a useful starting point if it tests the final deployed configuration, and asks for added elements and critical-error ceilings beyond aggregate accuracy."}, {"id": "0057", "p": "conditional", "g": "Endorses the examination-style approach but says benchmarks omit sensitivity to the claimed identity and seniority of the speaker, and should test populations differing in disease distribution."}, {"id": "0058", "p": "conditional", "g": "Endorses the examination-style approach but says benchmarks omit sensitivity to the claimed identity and seniority of the speaker, and should test populations differing in disease distribution."}, {"id": "0059", "p": "conditional", "g": "Benchmarks for calibration and subgroup performance should add an evidence-quality check, tested with historical-reversal cases, and treat subgroup evaluation as safety and usability rather than parity."}, {"id": "0061", "p": "conditional", "g": "Supports agentic element A.1 as drafted but says it misses step-report truthfulness, persistent state, absence claims and cross-session drift; proposes six added acceptance criteria."}, {"id": "0062", "p": "conditional", "g": "Expand safety benchmarking to assess confidence that valid input was received and treat fluent text generated from null sensor input as failure."}, {"id": "0064", "p": "conditional", "g": "Accepts benchmarking element S.3 on false confidence but asks that it cover confidence that any input was received, tested with null-input cases."}, {"id": "0067", "p": "conditional", "g": "Supports the benchmarking elements but wants a test-article configuration definition, a link from E.4 to human factors, and contamination controls with prespecified sponsor benchmark plans."}, {"id": "0074", "p": "conditional", "g": "Model-level competency benchmarking is important but should be complemented by system-level verification of the full sensor-to-output pathway, with benchmarks valued for credibility and traceability rather than size."}, {"id": "0076", "p": "conditional", "g": "Supports competency-based assessment but says the credentialing analogy cannot justify relaxed change control or shared accountability; evidence must be bound to a fixed configuration."}, {"id": "0078", "p": "conditional", "g": "Retain the element set, but specify session sampling position, score trajectories, add an input-integrity element, recategorize retrieved-content injection and automation bias under Safety, and allow device-independent records."}, {"id": "0082", "p": "conditional", "g": "Accepts benchmark gating but wants premise-aware items with human adjudication, age-stratified negatives, and fully disclosed sponsor-developed benchmarks instead of mandated independent ones."}, {"id": "0085", "p": "mixed", "g": "Offers no answer on construct validity; says benchmark evidence should be accompanied by structured, machine-checkable dataset documentation kept separate from performance evaluation."}, {"id": "0086", "p": "conditional", "g": "Says benchmarks alone are insufficient for systems influencing diagnosis, triage, treatment, escalation or monitoring."}, {"id": "0087", "p": "conditional", "g": "Finds competency-based assessment promising and agrees the deployed system should be evaluated, but says benchmarks must not become proxies for clinical validity."}, {"id": "0088", "p": "conditional", "g": "Calls benchmarking plus risk-proportionate confirmation the right frame but says it under-weights multi-turn information gathering and run-to-run reliability on the finished configuration."}, {"id": "0091", "p": "conditional", "g": "Considers the benchmarking elements and two-stage structure sound but says construct validity requires a six-component measurement-system qualification package plus a null comparison."}, {"id": "0093", "p": "conditional", "g": "Calls the competency approach useful and the element set well-constructed; asks for a provenance element, recognized external standards, benchmark sequestration and versioning, and architecture-specific applicability notes."}, {"id": "0094", "p": "conditional", "g": "Supports the competency-based approach if genuinely risk-proportionate; wants validity and utility separated per function and new elements for claims-evidence alignment and clinician usability."}, {"id": "0095", "p": "conditional", "g": "Supports the competency-based approach but asks that any benchmark gating a regulatory decision carry a documented reproducibility floor, since scores are single stochastic draws."}, {"id": "0097", "p": "conditional", "g": "Says benchmark examinations and retrospective datasets cannot reproduce clinical medicine and overall accuracy is insufficient; wants testing under realistic conditions including atypical and complex patients."}, {"id": "0098", "p": "conditional", "g": "Accepts the benchmarking structure but asks for added elements: fault-condition escalation tests, sycophancy resistance, three-state output classification with mandatory deferral, and verification-gate placement disclosure."}, {"id": "0099", "p": "support", "g": "Endorses the competency-based evaluation approach as directionally sound."}, {"id": "0100", "p": "conditional", "g": "Supports the competency-based structure and proposed elements, but wants architectural risk controls recognized in scoping benchmarking burden and sponsor-developed benchmarks accepted with validation."}], "themes": [{"name": "Add new benchmarking elements such as cybersecurity, provenance, input integrity", "ids": ["0005", "0010", "0018", "0025", "0028", "0032", "0035", "0045", "0049", "0051", "0054", "0056", "0062", "0078", "0093", "0094", "0098"], "n": 16}, {"name": "Require sequestered, contamination-controlled benchmark test sets", "ids": ["0003", "0005", "0010", "0013", "0017", "0035", "0045", "0051", "0056", "0067", "0093"], "n": 10}, {"name": "Require construct-validity evidence that benchmarks predict clinical performance", "ids": ["0003", "0005", "0017", "0028", "0048", "0049", "0051", "0091", "0093", "0100"], "n": 10}, {"name": "Benchmark the final deployed configuration, not the foundation model alone", "ids": ["0017", "0035", "0053", "0054", "0055", "0056", "0074", "0076", "0087"], "n": 9}, {"name": "Include adversarial, out-of-distribution, rare and messy test cases", "ids": ["0010", "0017", "0021", "0024", "0025", "0026", "0048", "0057", "0058", "0097"], "n": 9}, {"name": "Set separate pass/fail floors for high-consequence failures", "ids": ["0021", "0052", "0055", "0056"], "n": 4}, {"name": "Do not let benchmarks substitute for clinical evidence", "ids": ["0020", "0086", "0087"], "n": 3}]}, {"key": "clinical_confirmation_studies", "name": "Clinical confirmation", "qs": "Q11-12, Q14-15", "n": 41, "counts": {"conditional": 35, "support": 6}, "rows": [{"id": "0003", "p": "conditional", "g": "Calibrate to risk: shadow deployment suffices for HCP-supervised Class II; reserve RCTs for autonomous patient-facing Class III; use board-certified specialist, not median clinician, as comparator."}, {"id": "0005", "p": "conditional", "g": "Supports a confirmation ladder scaled to the highest-risk branch, with prespecified estimands, comparators tied to intended use and counterfactual care, and separate device-alone/team results."}, {"id": "0006", "p": "conditional", "g": "Supports the confirmation ladder; wants presumptive tiers keyed to risk, retrospective plus standardized patient evaluation for low-risk informational functions, and context-anchored comparators including usual information environment."}, {"id": "0009", "p": "conditional", "g": "Supports the range of confirmation approaches without mandatory prospective studies; asks for task-based adjudicator qualification, explicit blinding sequence, reported panel agreement, and prevailing-practice comparators."}, {"id": "0010", "p": "conditional", "g": "Supports graduated ladder if tier is driven by published risk mapping; shadow deployment default at intermediate risk; reject median-clinician standard; do not pool benchmark and confirmation estimates."}, {"id": "0012", "p": "support", "g": "Endorses risk-tiered clinical confirmation as part of the competency-based evaluation model, without further detail."}, {"id": "0013", "p": "conditional", "g": "Supports the graduated ladder, but wants tier selection driven by a published risk mapping, shadow deployment as default at intermediate risk, and no median-clinician standard."}, {"id": "0017", "p": "conditional", "g": "Confirmation rigor should match procedural consequence: retrospective for observational functions, prospective realistic evaluation for systems directing or executing invasive actions, with human-AI team measurement and absolute safety floor."}, {"id": "0018", "p": "conditional", "g": "Finds statistical and acceptance-criteria guidance insufficient; recommends practical statistical approaches, risk-adjusted thresholds, and human-AI team effectiveness measures."}, {"id": "0019", "p": "conditional", "g": "Insists on clinical trials rather than reduced confirmation requirements, with public transparency, informed consent, qualified clinicians, and a comparative device and pilot study."}, {"id": "0020", "p": "conditional", "g": "Confirmation hierarchy is a strength, but comparator selection and statistical standards are unresolved; evidence should escalate with risk up to prospective comparative trials."}, {"id": "0021", "p": "conditional", "g": "Median-clinician comparator understates model risk because misses replicate at scale; confirmation should compare downstream outcomes and workflow value against standard of care."}, {"id": "0024", "p": "conditional", "g": "Wants clinical validation to assess whether intended users can interpret and appropriately act on device output under realistic working conditions, supported by human factors evidence."}, {"id": "0025", "p": "conditional", "g": "Confirmation should grow more rigorous and prospective as systems become patient-facing, autonomous and consequential; comparators should include an evidence-based standard, current unaided practice, and the human-AI team."}, {"id": "0028", "p": "conditional", "g": "Supports risk-proportionate confirmation without mandatory prospective trials, requesting selection criteria, combined methods, justified comparators, and primary evaluation of human-AI teams."}, {"id": "0030", "p": "conditional", "g": "Supports proportionate clinical confirmation without mandatory prospective studies, but methods should address residual uncertainty as alternatives or complements, not a fixed hierarchy."}, {"id": "0035", "p": "support", "g": "Support clinical confirmation tailored to the actual decision pathway, using team performance for assisted care and autonomous performance when human review is ineffective."}, {"id": "0038", "p": "conditional", "g": "FDA's list of confirmation approaches is strong but incomplete; add hallucination/adversarial, bias, drift, multi-setting and human factors testing."}, {"id": "0043", "p": "conditional", "g": "Confirmation should follow a risk framework covering therapeutic area, exposure and real-world use; comparators should be meta-analytic data, a generalist-specialist mix, and 'no intervention'."}, {"id": "0045", "p": "conditional", "g": "Wants FDA's progression formalized as a five-tier evidence ladder set by risk, with clinically meaningful effect sizes and comparators reflecting the care patients would otherwise receive."}, {"id": "0046", "p": "conditional", "g": "Clinical confirmation should measure the human-AI team as actually deployed, while benchmarking measures the model alone."}, {"id": "0049", "p": "support", "g": "FDA's ladder of confirmation approaches is appropriate; selection should follow risk, reversibility, autonomy and reference standards, with comparators matched to the device's intended clinical role."}, {"id": "0051", "p": "conditional", "g": "Supports risk-scaled confirmation and absence-of-device comparators; wants formal transportability assessment for RWD and OUS data and no pooling of benchmark and confirmation evidence unless exchangeable."}, {"id": "0052", "p": "conditional", "g": "Agrees clinical confirmation is an important separate step; wants it scaled to criticality and residual uncertainty, temporal fidelity, separate reference standard and comparator, and human-AI team evaluation with non-inferiority."}, {"id": "0053", "p": "conditional", "g": "Clinical confirmation for physical-output functions should include physical evidence such as dimensional verification and mechanical testing, not only clinician adjudication; retrospective imaging evaluation is a natural first step."}, {"id": "0054", "p": "conditional", "g": "Where clinician supervision is part of intended use, the human-AI team may be the evaluation unit, but the practical effectiveness of oversight must itself be assessed."}, {"id": "0055", "p": "support", "g": "Confirmation method should scale with possible harm and remaining uncertainty, from retrospective to prospective; standard of care sets the minimum, and comparators should reflect care without the device."}, {"id": "0056", "p": "conditional", "g": "Comparators should match the task; specialist panels should set the reference for specialty functions, and both device-alone and human-device team performance should be evaluated."}, {"id": "0057", "p": "conditional", "g": "Wants sponsors to characterize anticipated deployment case mix and to provide human-AI team evidence of override and detection rates in representative workflows."}, {"id": "0058", "p": "conditional", "g": "Wants sponsors to characterize anticipated deployment case mix and to provide human-AI team evidence of override and detection rates in representative workflows."}, {"id": "0064", "p": "conditional", "g": "Keeps the Section V.C comparators but wants an absence-of-device comparator added as a pass/fail floor test, mandatory where the device is the only path to the output."}, {"id": "0070", "p": "conditional", "g": "Human-AI team performance should be the evaluation basis when an accountable clinician sees the output, provided the output's workflow position and ordering are fixed in the intended use."}, {"id": "0074", "p": "support", "g": "Clinical confirmation remains necessary where anatomy, physiology or real-world use conditions materially affect performance; controlled verification complements rather than replaces it."}, {"id": "0076", "p": "conditional", "g": "The comparator should be the standard of care, not the median clinician in practice, and human-AI team evaluation should measure automation bias."}, {"id": "0082", "p": "conditional", "g": "For functions intended for use when no professional is available, the primary comparator should be the unaided user, with standard of care second."}, {"id": "0086", "p": "conditional", "g": "Wants prospective clinical validation for high-risk systems, evaluation of the human-AI team including nurses, and comparison against safe, adequately staffed care."}, {"id": "0087", "p": "support", "g": "Says clinical confirmation remains essential with rigor proportional to risk, finds FDA's spectrum useful, and favors human-AI team evaluation for assistive uses."}, {"id": "0091", "p": "conditional", "g": "How far benchmarking can substitute for prospective study should depend on benchmark qualification; benchmark and confirmation evidence should not be naively pooled."}, {"id": "0093", "p": "conditional", "g": "Calls the confirmation ladder sound; wants selection justified on pre-declared dimensions, no pooling of benchmarking with confirmation evidence, and task-specific pre-specified comparators including counterfactual care."}, {"id": "0094", "p": "conditional", "g": "Endorses benchmarking followed by clinical confirmation but wants prospective studies reserved for the highest-risk band, longitudinal databanks listed as an approach, and human-AI team comparators for extender devices."}, {"id": "0095", "p": "conditional", "g": "Says statistically meaningful measurement requires comparing performance differences against the device's own run-to-run variation, reported per graded endpoint rather than pooled."}, {"id": "0097", "p": "conditional", "g": "Wants rigorous prospective clinical validation for systems that initiate or modify treatment, prescribe, decide on emergency evaluation, or make other high-consequence decisions."}, {"id": "0100", "p": "conditional", "g": "Supports risk-scaled clinical confirmation without mandatory prospective studies; asks for median-clinician default comparator, counterfactual comparators, and evaluation basis matched to deployed workflow."}], "themes": [{"name": "Evaluate human-AI team performance in realistic workflows", "ids": ["0005", "0017", "0018", "0024", "0028", "0046", "0052", "0054", "0056", "0057", "0058", "0070", "0086", "0094", "0100"], "n": 14}, {"name": "Use counterfactual care without the device as comparator", "ids": ["0005", "0020", "0025", "0043", "0045", "0051", "0052", "0064", "0082", "0100"], "n": 10}, {"name": "Prespecify acceptance criteria, margins and statistical precision", "ids": ["0005", "0010", "0013", "0020", "0045", "0056", "0076"], "n": 6}, {"name": "Reserve prospective studies for high-risk autonomous functions", "ids": ["0003", "0020", "0093", "0094", "0097"], "n": 5}, {"name": "Reject the median clinician as the comparator standard", "ids": ["0003", "0010", "0013", "0021", "0025", "0076"], "n": 5}, {"name": "Publish confirmation tiers keyed to the risk framework", "ids": ["0006", "0010", "0013", "0028", "0045"], "n": 4}, {"name": "Do not pool benchmarking and confirmation evidence into one estimate", "ids": ["0010", "0013", "0051", "0091", "0093"], "n": 4}]}, {"key": "synthetic_data", "name": "Synthetic data", "qs": "Q13", "n": 23, "counts": {"conditional": 20, "mixed": 3}, "rows": [{"id": "0003", "p": "conditional", "g": "Synthetic data useful for augmentation but creates circular validation risk; require out-of-distribution validation and real-world data as primary evidence for subpopulations."}, {"id": "0005", "p": "conditional", "g": "Synthetic data suits stress and rare scenarios but not standalone safety-critical claims; require independent real sentinels, lineage disclosure, and separate reporting."}, {"id": "0010", "p": "conditional", "g": "Synthetic data acceptable only for finding failures and rare-scenario stress testing; not for establishing performance rates; same-family generation barred for subgroup claims."}, {"id": "0013", "p": "conditional", "g": "Synthetic data is acceptable for finding failures and rare-event coverage but must not establish performance rates; same-family generation should be barred for subgroup claims."}, {"id": "0017", "p": "conditional", "g": "Synthetic data can supplement rare-event testing but must not substitute for representative real clinical data in high-consequence applications; report separately."}, {"id": "0020", "p": "conditional", "g": "Synthetic data is promising for rare pediatric presentations but risks circular blind spots; should augment, not replace, real-patient validation for substantial consequences."}, {"id": "0025", "p": "conditional", "g": "Synthetic data is useful for rare scenarios, privacy and stress testing, but synthetic patients are too tidy; it must model real human behavior and cannot replace it."}, {"id": "0028", "p": "conditional", "g": "Supports synthetic data as a supplement for rare cases and underrepresented subgroups, but warns of circularity and asks for validation, disclosure and distributional-shift analysis."}, {"id": "0035", "p": "conditional", "g": "Synthetic data may supplement but must not replace evidence from representative real distributions, especially for rare events and underrepresented groups."}, {"id": "0043", "p": "conditional", "g": "Synthetic subpopulations are feasible for any condition that is not rare or ultra-rare, provided the synthetic data derive from literature meta-analyses rather than single studies."}, {"id": "0045", "p": "conditional", "g": "Synthetic data is valuable for rare, adversarial and dangerous scenarios but must not become the main evidence; same-class generators can reproduce blind spots."}, {"id": "0048", "p": "conditional", "g": "Synthetic data can support schema testing, fault injection, rare cases and reproducibility but should not alone establish prevalence, subgroup or clinical performance, or benefit-risk."}, {"id": "0049", "p": "conditional", "g": "Synthetic data suits rare-event and stress testing but should complement real-world evidence; report it separately, document provenance, and guard against correlated blind spots from same-class generators."}, {"id": "0051", "p": "conditional", "g": "Under the statistical measurement and synthetic data questions, asks that subgroup performance claims rest on a minimum core of real data."}, {"id": "0052", "p": "conditional", "g": "Synthetic data are valuable for rare conditions, counterfactuals and difficult trajectories but should augment real data, have an explicit regulatory role, and not be pooled across purposes."}, {"id": "0055", "p": "conditional", "g": "Synthetic data suits scope, rare-hazard, adversarial and tool-failure tests but cannot replace real data on communication, behavior, workflow or underrepresented groups; important results need real-data confirmation."}, {"id": "0074", "p": "conditional", "g": "Simulated physiological signals are useful for repeatable robustness, boundary and regression testing, but only within the conditions the simulation is characterized to reproduce, with limitations stated."}, {"id": "0076", "p": "conditional", "g": "Synthetic data may supplement real data for stress tests and rare presentations but must come from a different model lineage and never solely support subgroup claims."}, {"id": "0085", "p": "mixed", "g": "Does not judge synthetic data; suggests an origin declaration distinguishing acquired, synthetic, simulated, phantom-derived and mixed sources, with the generation method recorded."}, {"id": "0091", "p": "conditional", "g": "Synthetic cases suit stress testing of rare, adversarial and safety-critical situations but cannot establish representativeness; labels must be verified independently of the generating model."}, {"id": "0093", "p": "conditional", "g": "Synthetic data suits rare-event, boundary, adversarial and multi-turn testing but is unreliable for underrepresented subgroups; acceptable only with safeguards and demonstrated representativeness."}, {"id": "0094", "p": "mixed", "g": "States that longitudinal real-world databanks outperform synthetic data for hard-to-sample populations and avoid the circularity risk of same-class synthetic generation."}, {"id": "0095", "p": "mixed", "g": "Under Question 13, says a null subgroup result is uninformative unless the sponsor states the minimum detectable difference for that analysis."}, {"id": "0096", "p": "conditional", "g": ""}], "themes": [{"name": "Disclose synthetic data generation method, provenance and lineage", "ids": ["0005", "0010", "0013", "0017", "0028", "0049", "0085", "0093"], "n": 7}, {"name": "Report synthetic and real-data results separately", "ids": ["0005", "0017", "0048", "0049", "0055", "0091"], "n": 6}, {"name": "Require generators independent of the device's model family", "ids": ["0010", "0013", "0045", "0049", "0055", "0076", "0093"], "n": 6}, {"name": "Do not use synthetic data alone for subgroup claims", "ids": ["0003", "0035", "0048", "0051", "0076"], "n": 5}, {"name": "Validate synthetic data against real-world distributions", "ids": ["0028", "0043", "0093"], "n": 3}, {"name": "Limit synthetic data to stress testing and rare scenarios", "ids": ["0010", "0013", "0048"], "n": 2}]}, {"key": "independent_third_party_testing", "name": "Third-party testing", "qs": "Q16", "n": 28, "counts": {"conditional": 22, "mixed": 2, "support": 3, "oppose": 1}, "rows": [{"id": "0003", "p": "conditional", "g": "Strongly supports third-party testing but warns of a testing duopoly; expand ASCA to AI benchmarking organizations with published fees and qualification criteria."}, {"id": "0005", "p": "conditional", "g": "Supports qualified third parties for sequestered datasets, red-teaming and adjudication, provided conflict rules, multiple providers, rotation and appeal mechanisms exist."}, {"id": "0006", "p": "conditional", "g": "Third parties are appropriate for sequestered benchmarks and adjudication panels, but fees and access must scale with sponsor size to avoid an incumbent toll gate."}, {"id": "0010", "p": "conditional", "g": "Supports third parties for benchmark administration, escrow and adjudication, with published recognition criteria, strict independence rules, and a preserved first-party pathway; notes capacity does not yet exist."}, {"id": "0013", "p": "conditional", "g": "Third parties suit benchmark administration, escrow and adjudication, but participation should be optional, with published recognition criteria, role separation and financial disclosure."}, {"id": "0017", "p": "conditional", "g": "Support independent participation in high-risk evaluation and audits, with financial independence, technical access, conflict safeguards, and freedom to report findings."}, {"id": "0019", "p": "mixed", "g": "Says an independent third party would be good and many bodies should be involved, but doubts such organizations' involvement would be possible; cites privacy limits."}, {"id": "0025", "p": "conditional", "g": "Supports third-party involvement, particularly for high-impact systems, provided independence is substantive and safeguards prevent cherry-picking and barriers to competition."}, {"id": "0028", "p": "conditional", "g": "Supports an expanded third-party role as adjudicators and stewards of sequestered benchmarks, piloted first and designed to avoid bottlenecks for smaller sponsors."}, {"id": "0029", "p": "support", "g": "Supports independent third parties in benchmarking, clinical confirmation, adjudication and standards, especially for sponsor-selected benchmarks, third-party models, and postmarket continuation decisions."}, {"id": "0035", "p": "conditional", "g": "Third parties can strengthen evaluation for sequestered assets, adversarial testing, adjudication, and cybersecurity, provided independence is disclosed and no exclusive certification structures limit competition."}, {"id": "0043", "p": "oppose", "g": "A third-party requirement would invite third-party organizations to game the system and turn testing into a profit center; he calls it inadvisable."}, {"id": "0045", "p": "conditional", "g": "Supports a network of FDA-overseen third parties building on ASCA and MDDT, provided there are multiple organizations, accreditation criteria, conflict rules and appeals."}, {"id": "0047", "p": "support", "g": "A credible framework should require independent premarket testing and recurring third-party audits."}, {"id": "0048", "p": "conditional", "g": "Says external adjudication, preregistration, blinded evaluation, locked tests or qualified reproduction may be appropriate when the sponsor controls test construction, tuning, scoring and interpretation."}, {"id": "0049", "p": "conditional", "g": "Third parties add value in benchmark custody, test-set maintenance, conformity testing and adjudication, provided sponsor accountability remains and no mandatory certification market limits competition."}, {"id": "0051", "p": "support", "g": "Sees a role for third parties in holding sequestered test sets independent of both sponsor and foundation model developer."}, {"id": "0055", "p": "conditional", "g": "Third parties can hold hidden test data, run benchmarks, review and audit, provided they are qualified, conflict-free, plural, and the manufacturer stays accountable."}, {"id": "0056", "p": "conditional", "g": "Supports independent third-party roles for moderate- and high-consequence functions, with specialty-matched adjudicators, defined independence criteria, and proportionality so it does not favor large sponsors."}, {"id": "0067", "p": "conditional", "g": "Third parties should qualify and maintain benchmark assets through the MDDT program rather than certify devices, with structural independence from sponsors and foundation model developers."}, {"id": "0071", "p": "conditional", "g": "Third-party involvement is structurally necessary for safety-critical elements S.1 to S.3, with independence defined at organizational and evaluator-model levels and open accreditation rather than a closed panel."}, {"id": "0076", "p": "conditional", "g": "Supports independent adjudicators and third-party sequestered datasets, with strict sponsor separation, validated independent LLM judges, and multiple accredited evaluators."}, {"id": "0085", "p": "mixed", "g": "Neither endorses nor rejects third-party roles; says standardized documentation would help sequestered-dataset custodians describe withheld assets consistently and is an early standards candidate."}, {"id": "0088", "p": "conditional", "g": "Favors shared item definitions and a hidden holdout, with licensed clinician review limited to decision gold."}, {"id": "0091", "p": "conditional", "g": "Prefers a hybrid of open, reproducible methods with independence targeted at reference-standard adjudication and sequestered safety-critical case sets, rather than a certification regime."}, {"id": "0093", "p": "conditional", "g": "Supports third-party participation modeled on ASCA and MDDT, with no exclusivity, non-discriminatory access, and final acceptance criteria and review decisions retained by CDRH."}, {"id": "0094", "p": "conditional", "g": "Accepts third-party involvement only as a voluntary, non-gating pathway like ASCA, with shared or subsidized infrastructure and recognition of CLIA/CAP accreditation."}, {"id": "0098", "p": "conditional", "g": "Agrees adjudicator independence should apply to LLM adjudicators but says organizational independence is insufficient; asks for disclosure of model-family relationship, decoding mode and training lineage."}, {"id": "0100", "p": "conditional", "g": "Supports third-party roles in sequestered datasets and adjudication, provided program design avoids incumbent moats, allows sponsor self-assessment, and sets independence criteria including for LLM adjudicators."}], "themes": [{"name": "Publish qualification and accreditation criteria for third-party evaluators", "ids": ["0003", "0005", "0010", "0013", "0028", "0045", "0049", "0056", "0071", "0100"], "n": 9}, {"name": "Require conflict-of-interest rules and financial relationship disclosure", "ids": ["0005", "0010", "0013", "0017", "0025", "0035", "0045", "0049", "0055", "0056"], "n": 9}, {"name": "Ensure multiple qualified providers; avoid exclusive certification markets", "ids": ["0003", "0005", "0035", "0045", "0049", "0055", "0076", "0093", "0100"], "n": 9}, {"name": "Build third-party roles on the ASCA and MDDT programs", "ids": ["0003", "0006", "0028", "0045", "0067", "0093"], "n": 6}, {"name": "Keep third-party testing voluntary; preserve a first-party pathway", "ids": ["0010", "0013", "0043", "0094", "0100"], "n": 4}, {"name": "Keep fees and access proportionate for small sponsors", "ids": ["0006", "0094", "0100"], "n": 3}, {"name": "Require LLM adjudicators independent of the device's model family", "ids": ["0071", "0076"], "n": 2}]}, {"key": "premarket_for_postmarket_tradeoff", "name": "Premarket-for-postmarket trade", "qs": "Q18", "n": 30, "counts": {"conditional": 28, "oppose": 2}, "rows": [{"id": "0003", "p": "conditional", "g": "Endorses the trade as sound and necessary, on condition that the postmarket plan is fully prespecified at authorization and legally enforceable with defined thresholds and triggers."}, {"id": "0005", "p": "conditional", "g": "Permit greater postmarket reliance only when detectability, monitoring latency and reversibility are proven; never for catastrophic, irreversible, or poorly observable risk."}, {"id": "0006", "p": "conditional", "g": "Supports accepting greater premarket uncertainty for postmarket monitoring, conditioned on prespecified thresholds, drift surveillance with independent clinician adjudication, transparent reporting, and rollback capability."}, {"id": "0008", "p": "conditional", "g": "Accepts the tradeoff only if monitoring programs demonstrate completeness of coverage and their own reliability before being credited toward reduced premarket evidence."}, {"id": "0010", "p": "conditional", "g": "Possibly right direction, but only with enforceable automatic conditions, demonstrated ability to observe endpoints, detection faster than harm, no irreversible harm, and excluded for high-consequence and agentic devices."}, {"id": "0013", "p": "conditional", "g": "May be the right direction, but only with enforceable conditions, observable endpoints, short detection latency and reversibility; unavailable for high-consequence, high-activity and agentic devices."}, {"id": "0017", "p": "conditional", "g": "Rejects the tradeoff for functions that direct or execute invasive actions or where harm precedes review; may be reasonable only for low-consequence, reversible, reviewable functions."}, {"id": "0020", "p": "conditional", "g": "Acceptable only for reversible, low-consequence functions; not for autonomous diagnosis, dosing, triage or other functions where failure first appears as patient harm."}, {"id": "0023", "p": "oppose", "g": "Rejects accepting greater premarket uncertainty; manufacturers must prove minimum accuracy before approval because postmarket review alone is insufficient for a new technology."}, {"id": "0024", "p": "conditional", "g": "Considers risk-proportional monitoring important but says it should not substitute for adequate premarket evidence in high-consequence applications."}, {"id": "0025", "p": "conditional", "g": "Sometimes acceptable within limits; tolerable premarket uncertainty should shrink as harm grows more serious, less reversible, more time-sensitive, and as systems become autonomous, patient-facing, or serve vulnerable populations."}, {"id": "0028", "p": "conditional", "g": "Supports the tradeoff if a specific, independently auditable monitoring plan is agreed premarket, with human oversight, sensitive metrics and prompt remediation; unsuitable for fully autonomous high-consequence functions."}, {"id": "0033", "p": "conditional", "g": "Trade-off can work only if the monitoring programme's detection capability is characterised, including affirmatively stated blind spots and time-to-detection, verified in operation."}, {"id": "0035", "p": "conditional", "g": "Reduced premarket evidence is acceptable only with prespecified, enforceable, actionable monitoring; inappropriate for high-consequence autonomous actions, rapid time-to-harm, weak reversibility, or inadequate interruption."}, {"id": "0045", "p": "conditional", "g": "Accepts the tradeoff for low- and moderate-risk systems with reversible harms and measurable performance; for high-consequence autonomous functions monitoring must supplement, not replace, premarket evidence."}, {"id": "0048", "p": "conditional", "g": "Monitoring should not substitute for premarket evidence; more residual uncertainty is acceptable only when failures are promptly detectable, consequences limited or reversible, exposure bounded and containment feasible."}, {"id": "0049", "p": "conditional", "g": "Acceptable when uncertainty is measurable, degradation is detectable before harm accumulates, rollback is feasible and evidence supports attribution; not for fully autonomous, severe, irreversible or poorly observable functions."}, {"id": "0051", "p": "conditional", "g": "Rebalancing is unsuitable for autonomous functions with severe consequences and for functions whose errors users cannot detect; conditions should be built into existing benefit-risk uncertainty guidance."}, {"id": "0052", "p": "conditional", "g": "Greater reliance on postmarket monitoring can be appropriate, but not as a general trade-off; only where residual uncertainty is observable, reversible and controllable, scaled to criticality."}, {"id": "0053", "p": "conditional", "g": "The trade is not appropriate for functions that produce permanent implants, because monitoring can detect but not remediate irreversible harm without surgery."}, {"id": "0054", "p": "conditional", "g": "Potentially appropriate for selected lower-risk functions with reversible outputs, effective human review, rapid rollback, comprehensive audit logging and limited autonomy; not presumed for irreversible, life-critical or high-consequence functions."}, {"id": "0055", "p": "conditional", "g": "Acceptable only when the manufacturer can detect and control remaining risk, with limited deployment, alert limits and rollback; unsuitable for irreversible, time-critical or undetectable-before-harm failures."}, {"id": "0075", "p": "conditional", "g": "Greater postmarket reliance may be appropriate with a supported initial benefit-risk determination, defined evidence targets, deadlines and exposure limits; not where serious harm could precede detection and intervention."}, {"id": "0076", "p": "conditional", "g": "A qualified no: reduced premarket evidence is acceptable only for non-directive, limited-consequence functions, and only once GenAI-specific postmarket infrastructure exists."}, {"id": "0077", "p": "conditional", "g": "Supports the trade in principle, but only as a failure-mode-specific exchange backed by demonstrated detection, independent records and timely notification, with four named device categories excluded."}, {"id": "0086", "p": "oppose", "g": "Rejects substituting postmarket monitoring for premarket evidence; monitoring must supplement strong premarket evidence, and patients should not become involuntary test subjects."}, {"id": "0090", "p": "conditional", "g": "Supports the shift in principle only if monitoring is continuous from first deployment, prespecified, independently verifiable, and has a consequence pathway; excludes high-autonomy severe-consequence functions and undetectable failures."}, {"id": "0093", "p": "conditional", "g": "Appropriate only under verifiable conditions: characterized monitoring, reversible or containable deployment, defined gates, lower or middle risk; not for irreversible high-consequence actions or unmonitorable settings."}, {"id": "0094", "p": "conditional", "g": "Wants the trade-off to be the default outside the highest-risk band, given a prespecified, resourced monitoring plan, rapid corrective action, and transparency allowing FDA audit."}, {"id": "0099", "p": "conditional", "g": "Supports exploring greater postmarket reliance in well-justified cases, provided monitoring targets the specific failure modes benchmarking left unresolved, with thresholds, owners and rollback criteria; not a substitute for premarket rigor."}, {"id": "0100", "p": "conditional", "g": "Strongly supports the tradeoff where sponsors precommit to a prespecified monitoring plan, triggers, independent clinician adjudication and corrective actions; less appropriate for autonomous high-consequence functions."}], "themes": [{"name": "Exclude high-risk or high-consequence functions", "ids": ["0005", "0010", "0013", "0017", "0020", "0024", "0025", "0028", "0035", "0045", "0048", "0049", "0051", "0052", "0054", "0075", "0076", "0090", "0093", "0094", "0100"], "n": 20}, {"name": "Require that failure be detectable before harm accumulates", "ids": ["0005", "0010", "0013", "0017", "0020", "0028", "0033", "0035", "0045", "0048", "0049", "0051", "0052", "0055", "0075", "0076", "0077", "0090", "0093"], "n": 18}, {"name": "Require that harm be reversible", "ids": ["0005", "0010", "0013", "0017", "0020", "0025", "0035", "0045", "0048", "0049", "0052", "0053", "0054", "0055", "0075", "0077", "0093"], "n": 16}, {"name": "Limit the trade to devices with restricted autonomy", "ids": ["0010", "0013", "0017", "0020", "0025", "0028", "0035", "0045", "0049", "0051", "0052", "0054", "0076", "0090", "0094", "0100"], "n": 15}, {"name": "Fix monitoring triggers and thresholds before authorization", "ids": ["0003", "0006", "0010", "0013", "0035", "0045", "0055", "0075", "0076", "0090", "0093", "0094", "0099", "0100"], "n": 13}, {"name": "Require proof that monitoring covers the failures it claims to catch", "ids": ["0005", "0008", "0010", "0013", "0033", "0035", "0049", "0054", "0076", "0077", "0090", "0093", "0099"], "n": 12}, {"name": "Require a rollback or stop mechanism", "ids": ["0006", "0028", "0035", "0049", "0052", "0054", "0055", "0090", "0093", "0094", "0099"], "n": 11}, {"name": "Make monitoring an enforceable condition of authorization", "ids": ["0003", "0010", "0013", "0028", "0035", "0075", "0100"], "n": 6}, {"name": "Require a monitor independent of the device model", "ids": ["0006", "0008", "0077", "0090", "0100"], "n": 5}, {"name": "Exclude users who cannot notice and report errors", "ids": ["0025", "0077", "0093"], "n": 3}]}, {"key": "postmarket_monitoring", "name": "Postmarket monitoring", "qs": "Q19-21", "n": 59, "counts": {"conditional": 52, "support": 5, "mixed": 2}, "rows": [{"id": "0003", "p": "conditional", "g": "Adopt risk-tiered cadence with defined triggers; set performance standards for supervisory agents before use; codify each stakeholder's monitoring role to prevent accountability diffusion."}, {"id": "0005", "p": "conditional", "g": "Supports the proposed approaches but wants event triggers, residual owners, thresholds, independently validated machine supervisors, and manufacturer-retained accountability."}, {"id": "0006", "p": "support", "g": "Endorses automated performance and drift surveillance with sampled adjudication by independent clinicians; crediting monitoring will push manufacturers to build instrumentation."}, {"id": "0008", "p": "conditional", "g": "Supports postmarket monitoring but wants documented coverage proportions and blind spots, plus degradation detection independent of the device's own reporting pathway."}, {"id": "0009", "p": "conditional", "g": "Asks FDA to define a role for health plans, which already audit assessments and hold outcome data, while the manufacturer keeps responsibility for the monitoring specification and thresholds."}, {"id": "0010", "p": "conditional", "g": "Three approaches sensible; cadence should be event-driven with calendar floor; supervisory agents must be architecturally independent, validated, and triage only; manufacturer accountability retained; MDR reportability must be resolved."}, {"id": "0012", "p": "support", "g": "Asks FDA to enable machine-based supervisory agents for postmarket monitoring of model drift."}, {"id": "0013", "p": "conditional", "g": "Finds the three approaches sensible; wants event-driven cadence with a calendar floor, independent validated supervisory agents, retained manufacturer accountability, and clarified MDR reportability."}, {"id": "0017", "p": "conditional", "g": "Surveillance must be version-, setting- and subgroup-specific, capturing near misses, overrides, failed takeovers and drift; cadence tied to risk triggers; supervisory agents must be validated."}, {"id": "0020", "p": "support", "g": "Rates postmarket surveillance a major strength; re-benchmarking, clinician review of real-world samples, and drift monitoring make sense for changing technologies."}, {"id": "0021", "p": "conditional", "g": "Sampled transcripts insufficient; monitor outcomes, re-benchmark after every upstream change, build pooled cross-manufacturer sentinel network, FDA device-model map, and require discontinuation plans."}, {"id": "0024", "p": "conditional", "g": "Supports postmarket monitoring and wants higher-risk devices to have a plan with data quality procedures, drift detection, thresholds, corrective measures and user communication of material changes."}, {"id": "0025", "p": "conditional", "g": "Monitoring should track patient consequences and near misses, not only model metrics; supervisory AI is useful if validated and human-governed; shared roles must not dilute accountability."}, {"id": "0027", "p": "conditional", "g": "Says sample-based clinician review and drift monitoring cannot show adversarial resistance; active adversarial re-benchmarking should be the principal evidence for Safety and Generalizability elements."}, {"id": "0028", "p": "conditional", "g": "Wants periodic re-benchmarking as the default, supplemented by drift detection; clinician review reserved for higher-risk devices or drift signals; supervisory agents need independent validation."}, {"id": "0029", "p": "conditional", "g": "Supports monitoring and re-benchmarking, while seeking formal continued-reliance requirements tied to assumptions, dependencies, thresholds, monitoring responsibilities, and reassessment actions."}, {"id": "0032", "p": "conditional", "g": "Wants monitoring broadened to outages, cybersecurity incidents, service lockouts and fleet-level events; supervisory agents must not be sole safeguard; manufacturer accountability must not be diffused; HTM stakeholders included."}, {"id": "0033", "p": "conditional", "g": "Supports sample-based clinician review as the right instinct but requires decision-rule criteria, measured adjudicator agreement, calibration, reconstructable records, accountable supervisory agents, and function-separated stakeholder roles."}, {"id": "0035", "p": "conditional", "g": "Reassessment should be periodic and event-triggered; supervisory agents must avoid circular assurance; ecosystem participants may contribute but manufacturers retain accountability; observability should be privacy-preserving."}, {"id": "0038", "p": "mixed", "g": "Under Q21, urges incorporating patients into governance, monitoring and evaluation and harmonizing with international regulators; does not evaluate specific monitoring methods."}, {"id": "0039", "p": "conditional", "g": "Postmarket monitoring should be proactive and cover drift, significant errors, near misses, clinician overrides, subgroup performance and real-world data changes."}, {"id": "0042", "p": "conditional", "g": "Degradation monitoring and supervisory agents work only if audit records distinguish human-authorized from agent-only actions; current HL7 FHIR standards cannot express this, and standards bodies should define attribution semantics."}, {"id": "0045", "p": "conditional", "g": "Recommends a continuous clinical-license model with scheduled, event-triggered and signal-triggered reassessment, validated supervisory agents, and primary manufacturer accountability within shared roles."}, {"id": "0046", "p": "conditional", "g": "The monitoring approaches and shared-ecosystem roles are well placed but work only if a portable, tamper-evident, privacy-preserving record of each human review travels with the determination."}, {"id": "0047", "p": "conditional", "g": "Supports continuous lifecycle monitoring but wants it mandatory and enforceable, with incident reporting, rollback procedures, penalties, and defined responsibilities so parties cannot shift accountability."}, {"id": "0048", "p": "conditional", "g": "A credible lifecycle program needs a versioned baseline, denominator and context data, indicators, investigation thresholds, incident linkage, ownership, and rollback or reassessment rules with defined triggers."}, {"id": "0049", "p": "conditional", "g": "FDA's monitoring approaches are useful but should be supplemented by reconstructable event-level records, hybrid cadence, independently evaluated supervisory agents and defined stakeholder control responsibilities."}, {"id": "0050", "p": "conditional", "g": "Institutions mostly cannot inventory or trace GenAI devices, so their role should be receiving and reporting, with manufacturers supplying version data and carrying the monitoring burden."}, {"id": "0051", "p": "conditional", "g": "Accepts the proposed approaches but wants a fourth, outcome-linked surveillance in claims and EHR data, with event-driven re-benchmarking cadence."}, {"id": "0052", "p": "conditional", "g": "Supports re-benchmarking, clinician review and degradation monitoring as complementary; wants monitoring of human-AI interaction and safeguards, periodic plus trigger-based reassessment, validated supervisory agents, and undiluted manufacturer accountability."}, {"id": "0053", "p": "conditional", "g": "On stakeholder roles, recommends a two-signature release separating surgeon approval from manufacturer release, keeping accountability with the manufacturer of record, with institutions and registries supporting outcome-linked surveillance."}, {"id": "0054", "p": "conditional", "g": "Monitoring should combine periodic review with event-triggered reassessment, tracking a list of failure signals and named triggers such as model updates, new tools and new populations."}, {"id": "0055", "p": "conditional", "g": "Monitoring should cover five areas and tie every signal to a specified action; machine supervisors must be tested as safety components; manufacturer stays accountable while others take defined roles."}, {"id": "0056", "p": "conditional", "g": "Wants periodic review combined with trigger-based reassessment, enriched sampling, clinically meaningful measures including near misses, and stakeholder contributions while manufacturers retain responsibility."}, {"id": "0057", "p": "conditional", "g": "Finds the proposed approaches reasonable but wants an examination-based model: fixed-cycle re-examination on the original competency set, also triggered by material change, not drift monitoring alone."}, {"id": "0058", "p": "conditional", "g": "Finds the proposed approaches reasonable but wants an examination-based model: fixed-cycle re-examination on the original competency set, also triggered by material change, not drift monitoring alone."}, {"id": "0060", "p": "conditional", "g": ""}, {"id": "0061", "p": "conditional", "g": "Accepts agentic devices in monitoring programs but warns that unreliable self-reports invalidate monitoring and users without independent knowledge cannot expose errors."}, {"id": "0062", "p": "conditional", "g": "Says postmarket mechanisms cannot see failures that leave no log, and supervisory agents must be validated against failure residue that mimics normal user behavior."}, {"id": "0063", "p": "conditional", "g": "Supports re-benchmarking, clinician review and degradation monitoring, but says periodic cadence misses within-session degradation; wants continuous verification against an independent record and event-based triggers."}, {"id": "0066", "p": "conditional", "g": "Supervisory agents can facilitate postmarket monitoring, but only if required to disclose evidence basis and age, fail on stale sources, and be validated without expert oversight."}, {"id": "0067", "p": "conditional", "g": "The three approaches are reasonable but lack action triggers; monitoring plans should prespecify parameters, sampling, alert and action limits and responses, modeled on Continued Process Verification."}, {"id": "0071", "p": "conditional", "g": "The three proposed approaches are necessary but individually insufficient; add continuous per-output verification, repeated-run reliability profiles, and independent governance of any supervisory agent."}, {"id": "0074", "p": "support", "g": "Citing Question 19, recommends that premarket benchmarks be designed for reuse as regression baselines so tests can be repeated under identical conditions after changes."}, {"id": "0075", "p": "conditional", "g": "Monitoring should tie each signal to an owner, intervention condition and response time; sampling should be independent of flagging; manufacturers keep responsibility; supervisory agents need role-specific evaluation."}, {"id": "0076", "p": "conditional", "g": "Postmarket surveillance is structurally weak for GenAI failure modes; shared ecosystem roles must not diffuse manufacturer accountability, and machine supervisors may triage but not replace human review."}, {"id": "0077", "p": "conditional", "g": "Monitoring must capture session-level failures, use independent records, and notify responsible people promptly; periodic reassessment and sampling can conceal degradation."}, {"id": "0081", "p": "conditional", "g": "Endorses shared stakeholder roles in oversight, including monitoring and reassessment, provided accountability is clearly defined and human oversight is meaningful."}, {"id": "0082", "p": "conditional", "g": "Supports periodic re-benchmarking as primary mechanism but says drift monitoring and real-session clinician review do not apply to frozen offline models, which need distinct obligations."}, {"id": "0083", "p": "conditional", "g": "Re-benchmarking, clinician review and degradation monitoring are acceptable only if manufacturers carry the surveillance burden and laboratory quality control stays risk-based, feasible and limited."}, {"id": "0085", "p": "mixed", "g": "Says periodic re-benchmarking comparisons are confounded unless the benchmark dataset's identity and version are explicitly documented."}, {"id": "0086", "p": "conditional", "g": "Calls monitoring essential; wants coverage of subgroups, adverse events, near misses and protected frontline reporting, no machine-only supervision, and manufacturers kept responsible."}, {"id": "0087", "p": "conditional", "g": "Supports periodic benchmarking, clinician review and degradation monitoring, provided monitoring stays proportional to risk and shared participation does not blur manufacturer, institution and clinician responsibilities."}, {"id": "0089", "p": "conditional", "g": "Postmarket surveillance of agentic devices is only tractable if each consequential action carries a retained, auditable, tamper-evident human authorization record."}, {"id": "0090", "p": "conditional", "g": "Endorses the three approaches as layers of one program, not a menu; wants event-driven cadence with a calendar floor and supervisory agents validated like devices."}, {"id": "0091", "p": "conditional", "g": "Change-triggered re-benchmarking should complement, not replace, periodic postmarket assessment, which remains necessary for undeclared drift; the adjudicator must be pinned in the baseline."}, {"id": "0092", "p": "conditional", "g": "Endorses periodic clinician review but wants random sampling combined with risk-based escalation, visible review states, role-scoped responsibility, and records supporting incident reconstruction."}, {"id": "0093", "p": "conditional", "g": "Views the three approaches as complementary and wants a risk-proportionate combination, an FDA trigger taxonomy, qualified supervisory agents, and decisions kept with the manufacturer."}, {"id": "0094", "p": "conditional", "g": "Supports the monitoring approaches, recommending laboratory quality-system cadences, independent clinician chart review, compensation-structure disclosure, and attention to deploying-site capacity."}, {"id": "0095", "p": "conditional", "g": "Supports the postmarket framework but says the reproducibility floor should set minimum detectable effect, sampling volume and triggers, and supervisory agents must not judge changes below their own instability."}, {"id": "0096", "p": "support", "g": "Supports proactive automated monitoring through recurring benchmarks, semantic drift analysis, and real-time guardrail and validator telemetry."}, {"id": "0097", "p": "conditional", "g": "Calls post-deployment surveillance essential; wants high-risk systems monitored for deterioration, unexpected behavior, errors and adverse outcomes, with significant AI-associated adverse events reportable and investigated."}, {"id": "0098", "p": "conditional", "g": "All proposed postmarket approaches depend on a retrievable decision record; persistent structured logging of safety-relevant decisions should be a precondition for postmarket monitoring credit."}, {"id": "0099", "p": "conditional", "g": "Supports monitoring and retained manufacturer accountability, requesting explicit evidence of effectiveness and monitoring tailored to unresolved failure modes."}, {"id": "0100", "p": "conditional", "g": "Confirms the proposed monitoring approaches are practical at scale; recommends stratified sampling, volume-scaled cadence, input-population drift monitoring, conditioned supervisory agents, and clarity that participating institutions are not manufacturers."}], "themes": [{"name": "Keep the manufacturer accountable despite shared ecosystem roles", "ids": ["0005", "0009", "0010", "0013", "0025", "0032", "0035", "0045", "0050", "0052", "0053", "0055", "0056", "0075", "0076", "0083", "0086", "0099"], "n": 17}, {"name": "Combine periodic reassessment with prespecified event triggers", "ids": ["0003", "0005", "0010", "0013", "0017", "0035", "0045", "0049", "0052", "0054", "0056", "0057", "0058", "0090", "0091", "0093"], "n": 14}, {"name": "Validate supervisory agents independently; keep human review in the loop", "ids": ["0005", "0010", "0013", "0017", "0025", "0028", "0035", "0049", "0052", "0062", "0066", "0071", "0075", "0076", "0086"], "n": 13}, {"name": "Monitor clinical outcomes, near misses and clinician overrides", "ids": ["0017", "0021", "0025", "0039", "0051", "0054", "0055", "0056", "0086", "0097"], "n": 10}, {"name": "Require reconstructable, tamper-evident records of device decisions", "ids": ["0025", "0033", "0046", "0047", "0049", "0089", "0092", "0098"], "n": 8}, {"name": "Prespecify monitoring thresholds, owners and required responses", "ids": ["0005", "0024", "0029", "0048", "0055", "0067", "0075", "0099"], "n": 8}, {"name": "Monitor performance by subgroup, version and setting", "ids": ["0017", "0039", "0086", "0087"], "n": 4}]}, {"key": "pccp_change_control", "name": "PCCP / change control", "qs": "Q22-24", "n": 49, "counts": {"conditional": 46, "support": 2, "mixed": 1}, "rows": [{"id": "0003", "p": "conditional", "g": "PCCPs must be substantially adapted via three change categories (locked, supervised, monitored); require 90-day notice and version pins for third-party foundation model changes."}, {"id": "0005", "p": "conditional", "g": "Classify changes by boundary delta, express PCCPs as an allowable change envelope, and require fingerprints, canary tests, hold and rollback for third-party model changes."}, {"id": "0006", "p": "conditional", "g": "Handle third-party model changes via version pinning, sponsor-side re-benchmarking gates before production, and a PCCP category for like-for-like model upgrades."}, {"id": "0008", "p": "conditional", "g": "Manufacturers' detection of third-party foundation model changes should not depend on signals mediated by the same system; requires an independent detection mechanism."}, {"id": "0010", "p": "conditional", "g": "Supports locked premarket baseline and tiered changes, but PCCP should lock invariants (escrowed benchmark suite, thresholds, protocol, rollback) rather than change lists; version pinning is a design control expectation."}, {"id": "0012", "p": "support", "g": "Wants flexible PCCPs to govern model drift so that supplemental clearances are not required."}, {"id": "0013", "p": "conditional", "g": "Supports a locked premarket baseline; proposes three change tiers, an invariant-based PCCP with escrowed benchmark suite, and version pinning as a design control expectation."}, {"id": "0016", "p": "conditional", "g": "Keep pre-specification but at higher abstraction; create a Performance-Bounded PCCP defining benchmarks, deviation bounds, monitoring, rollback, and cumulative impact tracking."}, {"id": "0017", "p": "conditional", "g": "Treat foundation-model changes as safety-relevant supply-chain events with contractual and technical controls; PCCPs should define bounded categories and rollback; authority-altering changes require FDA review."}, {"id": "0018", "p": "conditional", "g": "Third-party foundation model changes are the biggest practical challenge; wants locked versions, version traceability, supplier agreements, triggered revalidation, and PCCP provisions for model updates."}, {"id": "0021", "p": "conditional", "g": "Sponsors must hold enforceable information rights with model providers; inability to evaluate upstream dependency is itself an unresolved device risk."}, {"id": "0024", "p": "conditional", "g": "Asks FDA to address foundation model updates specifically and to require PCCPs to define modification bounds, acceptance standards, verification protocol, monitoring period, rollback and user notification."}, {"id": "0025", "p": "conditional", "g": "Re-testing should scale with whether a change can alter safety-critical behavior; define safety invariants for unpredictable changes; manufacturers stay responsible for third-party model changes."}, {"id": "0027", "p": "conditional", "g": "Wants compression named as a change category, validated on the exact deployed artifact, and not presumed low-impact for QMS or PCCP treatment; fine-tuning not presumed risk-reducing."}, {"id": "0029", "p": "conditional", "g": "Continued reliance should depend on evidence that authorization assumptions remain valid; higher-risk sponsors should prespecify material-change triggers, dependencies, and revalidation actions."}, {"id": "0030", "p": "conditional", "g": "Third-party foundation model changes are significant but not novel; build on SOUP/IEC 62304 principles with device-level controls, not full visibility, plus behavioral change detection."}, {"id": "0032", "p": "conditional", "g": "Manufacturers should stay accountable for third-party model changes; contracts, version pinning, staged deployment, rollback and notification belong in the safety case, and provider withdrawal must also be planned for."}, {"id": "0035", "p": "conditional", "g": "Change control should classify changes by impact pathway, not component type; PCCPs can define envelopes and triggers; third-party model changes need contractual and technical detection controls."}, {"id": "0039", "p": "conditional", "g": "Supports re-benchmarking against the premarket baseline; wants defined risk-based reassessment triggers for updates, retraining and third-party foundation-model changes."}, {"id": "0045", "p": "conditional", "g": "Wants modifications tiered by clinical impact, PCCPs that define boundaries and revalidation triggers rather than specific changes, and contractual and technical guardrails for third-party model updates."}, {"id": "0047", "p": "conditional", "g": "Every material change, including foundation model changes, should be documented and assessed, and independently revalidated before broader deployment when it could affect clinical performance or safety."}, {"id": "0048", "p": "conditional", "g": "Manufacturers using third-party foundation models should keep a versioned dependency inventory with change notice, impact assessment and rollback or pinning; uncontrolled dependencies should narrow permitted use."}, {"id": "0049", "p": "conditional", "g": "Reassessment should be proportional using a change taxonomy; PCCPs can prespecify bounded change classes and governance processes; third-party model changes need version attestation, notification, pinning and fail-safe responses."}, {"id": "0050", "p": "conditional", "g": "Handling third-party model changes requires timely institutional notification, including information about downstream effects and clearly specified communication channels."}, {"id": "0051", "p": "conditional", "g": "Proposes adapting PCCPs from prespecifying modifications to prespecifying the verification protocol, a change-envelope fixing baseline, re-benchmarking protocol, acceptance criteria and intended-use boundaries."}, {"id": "0052", "p": "conditional", "g": "Extending PCCPs to GenAI is reasonable, using procedural and boundary-based prespecification, staged autonomy with downgrading, cumulative-change review, foundation-model regression testing, requalification, and case-level reconstructability."}, {"id": "0053", "p": "conditional", "g": "PCCP concepts can work if plans enumerate interaction triggers across model and manufacturing changes, with tiered re-benchmarking and pinned third-party model versions."}, {"id": "0054", "p": "conditional", "g": "Change control should respond to behaviour change across the full system, including third-party model updates without code change, using a controlled model registry and staged promotion."}, {"id": "0055", "p": "conditional", "g": "Use quality systems for nonbehavioral changes and bounded PCCPs for specified changes; require submissions outside limits and effective control of foundation-model updates."}, {"id": "0057", "p": "conditional", "g": "Accepts impact-scaled re-benchmarking for proficiency elements but wants safety and robustness elements re-run in full after any model or guardrail change; PCCPs should prespecify the re-examination."}, {"id": "0058", "p": "conditional", "g": "Accepts impact-scaled re-benchmarking for proficiency elements but wants safety and robustness elements re-run in full after any model or guardrail change; PCCPs should prespecify the re-examination."}, {"id": "0062", "p": "conditional", "g": "Says Question 24 wrongly assumes manufacturers know which model is running; component identification must precede change detection, with version pinning required and PCCPs limited to pinned versions."}, {"id": "0067", "p": "conditional", "g": "PCCPs parallel ICH Q12 change management protocols; change categories should be defined by which benchmarking elements a change could affect, mapped prospectively in the PCCP."}, {"id": "0071", "p": "conditional", "g": "Modifications should be tiered by reversibility and blast radius; low-tier changes can sit in a PCCP, while changes touching safety-critical behaviors need re-benchmarking before deployment."}, {"id": "0074", "p": "support", "g": "Re-testing scope after a change should be risk-based: targeted regression for limited modifications, broader re-verification for changes affecting sensing, processing, algorithm behavior or clinical output."}, {"id": "0075", "p": "conditional", "g": "Supports impact-based change assessment across the full deployed configuration using the PCCP structure where applicable, with documented supplier guarantees, explicit adoption decisions and tested fallbacks for model updates."}, {"id": "0076", "p": "conditional", "g": "PCCPs suit sponsor-controlled changes such as prompts or guardrails, but third-party model changes need version pinning, change notification, no silent updates and re-benchmarking."}, {"id": "0082", "p": "conditional", "g": "Argues pinned, hash-verified on-device weights eliminate third-party-initiated change; fine-tuned open-weight models on a fixed base version should count as manufacturer-controlled."}, {"id": "0083", "p": "conditional", "g": "Manufacturers, not laboratories, should detect third-party foundation-model changes, evaluate impact, re-benchmark the finished device and communicate required actions to laboratories."}, {"id": "0085", "p": "mixed", "g": "Says change control that relies on a premarket baseline requires the baseline evaluation asset's identity to be verifiable, not assumed."}, {"id": "0086", "p": "conditional", "g": "Wants outputs traceable to a model version, material changes detected before affecting care with revalidation and FDA review, and suspension and rollback controls for health systems."}, {"id": "0087", "p": "conditional", "g": "Supports risk-based reassessment of changes rather than automatic full revalidation, and wants PCCPs to describe categories, boundaries and acceptance criteria instead of predicting every modification."}, {"id": "0088", "p": "conditional", "g": "Requests repeat evaluation of finished configurations whenever foundation-model identifiers or prompt templates change, with guardrails tested across turns."}, {"id": "0090", "p": "conditional", "g": "Says current PCCP categories fit discrete retraining and map poorly to GenAI; wants categories by system component, mapped re-benchmarking scope, and layered detection of third-party model changes."}, {"id": "0091", "p": "conditional", "g": "Re-benchmarking should scale to modification type via a published mapping, with an invariant safety regression core and bridge testing when the automated adjudicator changes."}, {"id": "0092", "p": "conditional", "g": "Supports proportionate reevaluation, triggered when changes materially alter how the device receives information, exercises authority, or affects clinical action, such as new models, tools or agent capabilities."}, {"id": "0093", "p": "conditional", "g": "Proposes a three-tier change taxonomy and shifting PCCPs from predetermined changes to predetermined gates, with layered contractual, technical and procedural controls for third-party model updates."}, {"id": "0094", "p": "conditional", "g": "Endorses the existing PCCP framework and asks that it be extended to carry the prespecified postmarket monitoring plan rather than creating a parallel mechanism."}, {"id": "0096", "p": "conditional", "g": "Third-party model updates are manageable if manufacturers combine local model locking, QMS multi-version lifecycle controls and automated drift detection, with FDA guidance recognizing these."}, {"id": "0099", "p": "conditional", "g": "Supports recognising non-code change sources, and asks that prompt, retrieval, guardrail and orchestration changes affecting output or risk controls be explicitly treated as device changes under change control."}, {"id": "0100", "p": "conditional", "g": "Supports PCCPs for bounded foundation-model updates passing regression gates, with architectural containment, version pinning, contractual notice, and shadow evaluation."}], "themes": [{"name": "Require pinned, identifiable third-party model versions", "ids": ["0003", "0006", "0010", "0013", "0018", "0021", "0032", "0035", "0048", "0049", "0053", "0062", "0076", "0090", "0100"], "n": 14}, {"name": "Reframe PCCPs around change envelopes and acceptance criteria", "ids": ["0005", "0010", "0013", "0016", "0017", "0024", "0035", "0045", "0049", "0051", "0087", "0093"], "n": 11}, {"name": "Require rollback capability for failed changes", "ids": ["0005", "0016", "0017", "0024", "0025", "0032", "0035", "0045", "0048", "0055", "0086"], "n": 11}, {"name": "Require contractual advance notice of foundation model changes", "ids": ["0003", "0010", "0017", "0021", "0035", "0076", "0090", "0100"], "n": 8}, {"name": "Adopt a tiered change taxonomy scaled to clinical impact", "ids": ["0003", "0013", "0045", "0049", "0053", "0071", "0093"], "n": 7}, {"name": "Require technical detection of unannounced model changes", "ids": ["0005", "0008", "0013", "0030", "0035", "0090"], "n": 6}, {"name": "Re-benchmark before any model change reaches clinical use", "ids": ["0006", "0076", "0093", "0100"], "n": 4}]}, {"key": "foundation_model_maf", "name": "Foundation model MAF", "qs": "Q25", "n": 25, "counts": {"conditional": 22, "support": 2, "oppose": 1}, "rows": [{"id": "0003", "p": "conditional", "g": "MAF concept sound but voluntary participation will fail; impose higher premarket evidence burden on devices lacking adequate MAFs and specify minimum MAF content."}, {"id": "0005", "p": "conditional", "g": "MAFs can help if current, versioned and safety-relevant, but voluntary gaps must remain explicit and sponsors still need independent dependency assurance."}, {"id": "0006", "p": "conditional", "g": "Advance voluntary MAFs with update-notification commitments and healthcare evaluation summaries, and signal more predictable review for devices on MAF-holding models to create developer incentive."}, {"id": "0008", "p": "conditional", "g": "Voluntary MAFs share developers' self-detection limitations; filers should be required to disclose their internal detection and monitoring methodology as a condition of reference."}, {"id": "0010", "p": "conditional", "g": "Create the voluntary MAF but pair it with a manufacturer-side default that absent an MAF the manufacturer bears full model characterization burden; MAF must be current and not substitute for device evaluation."}, {"id": "0012", "p": "support", "g": "Wants the Foundation Model MAF program fully operationalized so downstream developers can reference upstream model information without exposure of proprietary weights."}, {"id": "0013", "p": "conditional", "g": "Recommends creating the voluntary MAF, paired with a default that manufacturers bear the full model characterization burden when no MAF exists, creating procurement pressure."}, {"id": "0017", "p": "conditional", "g": "MAFs would help if current and technically meaningful, but voluntary participation may be insufficient; consider expecting MAF or equivalent when sponsors rely materially on third-party models."}, {"id": "0021", "p": "conditional", "g": "A voluntary master file alone is insufficient; enforceable information rights with model providers are needed."}, {"id": "0025", "p": "conditional", "g": "MAFs could improve efficiency and reduce duplication, but voluntary transparency must not replace required safety information or diffuse downstream manufacturer accountability."}, {"id": "0027", "p": "conditional", "g": "Says a MAF describes the developer's model, not the fine-tuned, compressed, configured deployed artifact, so referencing submissions should document downstream transformations."}, {"id": "0032", "p": "conditional", "g": "Master files could also carry support and deprecation policies, version pinning, incident notification, continuity and provider-exit terms, with the device sponsor remaining responsible."}, {"id": "0035", "p": "conditional", "g": "A voluntary MAF is useful if versioned, updateable, and contractually referenced, with tiered disclosure; it should feed a sponsor-maintained AI-SBOM rather than replace it."}, {"id": "0045", "p": "conditional", "g": "Calls the MAF a good concept but insufficient alone for highest-risk systems; wants it paired with contractual disclosure requirements and FDA authority to request more."}, {"id": "0049", "p": "conditional", "g": "Voluntary foundation model master files could reduce duplicated review if kept current with safety-relevant content, but must not replace sponsor validation of the final configured device."}, {"id": "0050", "p": "conditional", "g": "Accepts Master Files but says confidential-only files foreclose independent scrutiny; recommends a tiered structure with a narrow public subset, expanded only after evaluating participation."}, {"id": "0054", "p": "support", "g": "Supports voluntary foundation model master files, lists useful content, and says a file should not amount to FDA approval; responsibility stays with the downstream manufacturer."}, {"id": "0055", "p": "conditional", "g": "A MAF gives useful dependency information and reduces repeated work but cannot prove final device safety; CDRH needs another method when providers give insufficient information."}, {"id": "0057", "p": "conditional", "g": "If MAFs proceed, content should include refusal behavior, content-policy changes between versions and output stability under user framing; referencing MAFs in re-examination could incentivize developers."}, {"id": "0058", "p": "conditional", "g": "If MAFs proceed, content should include refusal behavior, content-policy changes between versions and output stability under user framing; referencing MAFs in re-examination could incentivize developers."}, {"id": "0076", "p": "oppose", "g": "A voluntary program will draw disclosure only from developers with little to hide; disclosure should be mandatory or the sponsor should carry the full evidence burden."}, {"id": "0079", "p": "conditional", "g": "Supports master files for stable content such as model family and change commitments, but says documents cannot track minute-scale model substitution; wants continuous provenance records instead."}, {"id": "0087", "p": "conditional", "g": "Supports exploring voluntary Foundation Model MAFs, provided a MAF is not FDA authorization of the model and does not shift responsibility from the device manufacturer."}, {"id": "0093", "p": "conditional", "g": "Supports voluntary MAFs as the lightest-touch mechanism; asks for a structured template and a sponsor-side documentation floor where no MAF is filed."}, {"id": "0094", "p": "conditional", "g": "Asks that any Foundation Model MAF program remain strictly voluntary and non-gating, with shared infrastructure available to small entities."}, {"id": "0099", "p": "conditional", "g": "A voluntary Foundation Model MAF could be useful for information-sharing, provided reliance on it does not shift the sponsor's responsibility for detecting and controlling foundation model changes."}, {"id": "0100", "p": "conditional", "g": "Supports voluntary Foundation Model MAFs provided absence of an MAF creates no presumption against a device and cloud platform intermediaries can be MAF holders."}], "themes": [{"name": "Specify minimum MAF content: versions, limitations, evaluations, update policy", "ids": ["0003", "0005", "0006", "0010", "0013", "0017", "0032", "0035", "0045", "0049", "0055", "0057", "0058", "0087", "0093"], "n": 13}, {"name": "Keep sponsor responsible for final device despite MAF reliance", "ids": ["0005", "0010", "0013", "0025", "0032", "0049", "0055", "0087", "0099"], "n": 8}, {"name": "Supplement voluntary MAFs with mandatory or contractual disclosure", "ids": ["0017", "0021", "0045", "0076"], "n": 4}, {"name": "Require MAF updates and change notification when models change", "ids": ["0006", "0010", "0013", "0025"], "n": 3}, {"name": "Place full characterization burden on sponsors lacking an MAF", "ids": ["0010", "0013", "0076"], "n": 2}, {"name": "Keep MAFs voluntary with no presumption against non-participants", "ids": ["0094", "0100"], "n": 2}]}], "subs": {"0003": {"who": "Walnut Hill Medical (WHM)", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0003", "type": "consultancy_or_law_firm"}, "0005": {"who": "Alfred T. McBride", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0005", "type": "other"}, "0006": {"who": "BellyMD, Inc.", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0006", "type": "device_or_digital_health_company"}, "0009": {"who": "Cara AI, Inc.", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0009", "type": "device_or_digital_health_company"}, "0010": {"who": "Hari Prakash Chanumolu", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0010", "type": "individual_regulatory_quality_pro"}, "0012": {"who": "Anonymous", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0012", "type": "unknown"}, "0013": {"who": "Hari Prakash Chanumolu", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0013", "type": "individual_regulatory_quality_pro"}, "0015": {"who": "Nathan Sabich", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0015", "type": "individual_engineer_data_scientist"}, "0017": {"who": "Douglas Stoddard", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0017", "type": "individual_clinician"}, "0019": {"who": "Shannon Nadia Kamalaker", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0019", "type": "individual_clinician"}, "0020": {"who": "Mark Simonian MD FAAP", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0020", "type": "individual_clinician"}, "0021": {"who": "Brainstorm: The Stanford Lab for Mental Health Innovation, Stanford University School of Medicine", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0021", "type": "academic_or_research_institution"}, "0024": {"who": "Srividya Narayanan", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0024", "type": "individual_regulatory_quality_pro"}, "0025": {"who": "Nurse Deb Speaking, Media & Consulting, LLC", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0025", "type": "consultancy_or_law_firm"}, "0027": {"who": "SichGate Inc.", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0027", "type": "device_or_digital_health_company"}, "0028": {"who": "Aruna Badiga", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0028", "type": "individual_regulatory_quality_pro"}, "0029": {"who": "Persistence Analytics Group LLC", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0029", "type": "consultancy_or_law_firm"}, "0030": {"who": "QRx Partners", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0030", "type": "consultancy_or_law_firm"}, "0032": {"who": "Gregory J. Marcisz", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0032", "type": "individual_engineer_data_scientist"}, "0035": {"who": "VivaSecuris", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0035", "type": "device_or_digital_health_company"}, "0037": {"who": "Vizma Carver", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0037", "type": "individual_engineer_data_scientist"}, "0038": {"who": "Mitchell Berger", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0038", "type": "individual_clinician"}, "0039": {"who": "Comment from Anonymous", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0039", "type": "unknown"}, "0042": {"who": "Behavioral Health Open Source (bh-healthcare.org)", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0042", "type": "individual_engineer_data_scientist"}, "0043": {"who": "Ben Locwin", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0043", "type": "individual_clinician"}, "0045": {"who": "Ravi R. Pankhaniya", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0045", "type": "individual_clinician"}, "0046": {"who": "Blaine Warkentine", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0046", "type": "individual_clinician"}, "0047": {"who": "QC Healthcare", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0047", "type": "other"}, "0048": {"who": "VitaSignal", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0048", "type": "device_or_digital_health_company"}, "0049": {"who": "OneSource Solutions International (OSSI)", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0049", "type": "device_or_digital_health_company"}, "0051": {"who": "Sehouenou Alberic Candide Ahouehome", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0051", "type": "individual_academic"}, "0052": {"who": "Furtwangen University", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0052", "type": "individual_academic"}, "0053": {"who": "Krishna Sai Koka", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0053", "type": "individual_academic"}, "0054": {"who": "Sitora Healthcare Digital", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0054", "type": "device_or_digital_health_company"}, "0055": {"who": "Newton’s Tree", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0055", "type": "device_or_digital_health_company"}, "0056": {"who": "Manuj Agarwal", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0056", "type": "individual_clinician"}, "0057": {"who": "DrEthan AI Aesthetics", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0057", "type": "individual_clinician"}, "0058": {"who": "DrEthan AI Aesthetics", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0058", "type": "individual_clinician"}, "0065": {"who": "The Christman AI Project / Luma Cognify AI. LLC", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0065", "type": "device_or_digital_health_company"}, "0067": {"who": "Bhasker Sambar", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0067", "type": "individual_regulatory_quality_pro"}, "0076": {"who": "Navid Farr", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0076", "type": "individual_regulatory_quality_pro"}, "0080": {"who": "Oura", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0080", "type": "device_or_digital_health_company"}, "0082": {"who": "Red Kit", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0082", "type": "device_or_digital_health_company"}, "0086": {"who": "Michelle Bernabe", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0086", "type": "individual_clinician"}, "0087": {"who": "S. Joseph Sirintrapun", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0087", "type": "individual_clinician"}, "0088": {"who": "Qiong Liu", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0088", "type": "individual_engineer_data_scientist"}, "0089": {"who": "Orzyn", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0089", "type": "individual_clinician"}, "0091": {"who": "Conefia LLC", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0091", "type": "device_or_digital_health_company"}, "0093": {"who": "Steven Zhao", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0093", "type": "individual_regulatory_quality_pro"}, "0094": {"who": "Profound Ventures Corporation and Guidance Global Consulting", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0094", "type": "device_or_digital_health_company"}, "0097": {"who": "Stephanie Lewis, M.D.", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0097", "type": "individual_clinician"}, "0098": {"who": "Royal College of Surgeons Ireland; Vox / VoxMedical", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0098", "type": "device_or_digital_health_company"}, "0099": {"who": "Matthew Collins", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0099", "type": "individual_regulatory_quality_pro"}, "0100": {"who": "Clearstep Inc.", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0100", "type": "device_or_digital_health_company"}, "0023": {"who": "Comment from Anonymous", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0023", "type": "patient_or_member_of_public"}, "0026": {"who": "Comment from Rama Doddi", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0026", "type": "individual_regulatory_quality_pro"}, "0061": {"who": "The Christman AI Project / Luma Cognify AI", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0061", "type": "device_or_digital_health_company"}, "0084": {"who": "Comment from Jenna Chuan", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0084", "type": "patient_or_member_of_public"}, "0018": {"who": "Jagadeesha Pampapathi", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0018", "type": "individual_engineer_data_scientist"}, "0059": {"who": "Middle Wave LLC / Teach & Serve", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0059", "type": "individual_clinician"}, "0062": {"who": "The Christman AI Project / Luma Cognify AI. LLC", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0062", "type": "device_or_digital_health_company"}, "0064": {"who": "The Christman AI Project / Luma Cognify AI", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0064", "type": "device_or_digital_health_company"}, "0074": {"who": "WhaleTeq Co., Ltd.", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0074", "type": "device_or_digital_health_company"}, "0078": {"who": "The Christman AI Project and Robotics Division", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0078", "type": "device_or_digital_health_company"}, "0085": {"who": "Princeton Medical Systems", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0085", "type": "test_lab_or_standards_body"}, "0095": {"who": "Rohith Reddy Bellibatlu", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0095", "type": "individual_academic"}, "0070": {"who": "Pallas Kliniken", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0070", "type": "individual_clinician"}, "0096": {"who": "Comment from Saurabh Sharma", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0096", "type": "unknown"}, "0071": {"who": "Orinyx", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0071", "type": "device_or_digital_health_company"}, "0008": {"who": "Supernova Technologies", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0008", "type": "other"}, "0033": {"who": "Shara Gospel", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0033", "type": "individual_regulatory_quality_pro"}, "0075": {"who": "Brandon Kaplan", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0075", "type": "individual_engineer_data_scientist"}, "0077": {"who": "The Christman AI Project and Robotics Division", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0077", "type": "device_or_digital_health_company"}, "0090": {"who": "Sentir Health Inc.", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0090", "type": "device_or_digital_health_company"}, "0050": {"who": "Chirag Kanitkar", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0050", "type": "other"}, "0060": {"who": "Comment from Christopher Tobias", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0060", "type": "unknown"}, "0063": {"who": "The Christman AI Project / Luma Cognify AI. LLC", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0063", "type": "device_or_digital_health_company"}, "0066": {"who": "The Christman AI Project / Luma Cognify AI. LLC", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0066", "type": "device_or_digital_health_company"}, "0081": {"who": "Comment from Cory Stephens", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0081", "type": "unknown"}, "0083": {"who": "Sihem Khelifa", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0083", "type": "individual_clinician"}, "0092": {"who": "Xiangyu Guo", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0092", "type": "individual_academic"}, "0016": {"who": "Comment from Nathan Sabich", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0016", "type": "unknown"}, "0079": {"who": "The Christman AI Project and Robotics Division", "url": "https://www.regulations.gov/comment/FDA-2026-N-7874-0079", "type": "device_or_digital_health_company"}}};
  var PL = {oppose:'Oppose',mixed:'Mixed',conditional:'With changes',support:'Support'};
  var $ = function(id){return document.getElementById(id);};
  var esc = function(s){return String(s).replace(/[&<>"]/g,function(c){return {'&':'&amp;','<':'&lt;','>':'&gt;','"':'&quot;'}[c];});};
  var sel = $('g18-sel'), p = null, theme = null;
  D.proposals.forEach(function(x){var o=document.createElement('option');o.value=x.key;o.textContent=x.name+' ('+x.qs+')';sel.appendChild(o);});
  $('g18-meta').textContent = 'Data as of '+D.generated+'. Themes are grouped by a language model from my paraphrases and can overlap. Follow a link for the original text on Regulations.gov.';
  function pick(){
    p = D.proposals.filter(function(x){return x.key===sel.value;})[0]; theme = null;
    var asking = (p.counts.conditional||0)+(p.counts.oppose||0)+(p.counts.mixed||0);
    $('g18-sub').textContent = p.n+' submitters address this proposal; '+asking+' ask for changes or oppose it. Bars count the submitters behind each recurring request. Select one to see who.';
    var max = Math.max.apply(null,p.themes.map(function(t){return t.n;}).concat([1]));
    $('g18-bars').innerHTML='';
    p.themes.forEach(function(t,i){var bt=document.createElement('button');bt.type='button';bt.className='q-bar';bt.id='g18-t-'+i;bt.dataset.i=i;bt.innerHTML='<span>'+esc(t.name)+'</span><span class="t"><i style="width:'+(100*t.n/max)+'%"></i></span><span class="n">'+t.n+'</span>';bt.addEventListener('click',function(){theme=theme===i?null:i;draw();});$('g18-bars').appendChild(bt);});
    draw();
  }
  function draw(){
    var ids = theme===null ? null : p.themes[theme].ids;
    var rows = p.rows.filter(function(r){return !ids || ids.indexOf(r.id)>-1;});
    $('g18-body').innerHTML = rows.map(function(r){var s=D.subs[r.id];return '<tr><td><a href="'+esc(s.url)+'" target="_blank" rel="noopener">'+esc(s.who)+'</a></td><td><span class="q-pill'+(r.p==='oppose'?' no':r.p==='support'?' yes':'')+'">'+PL[r.p]+'</span></td><td>'+esc(r.g)+'</td></tr>';}).join('');
    $('g18-tcap').textContent = theme===null ? 'All comments on this proposal ('+rows.length+')' : p.themes[theme].name+' ('+rows.length+' comments)';
    $('g18-clear').hidden = theme===null;
    Array.prototype.forEach.call(document.querySelectorAll('#g18-bars .q-bar'),function(x){x.setAttribute('aria-pressed',String(+x.dataset.i===theme));});
  }
  sel.addEventListener('change',pick); $('g18-clear').addEventListener('click',function(){theme=null;draw();});
  pick();
})();
</script>

#### Scope and risk (Questions 1 to 6)

The two-axis risk framework draws 52 submitters, and 47 want it changed. Twenty-three ask FDA to add reversibility of harm and time to harm as risk modifiers, 15 ask for traceability of an output to its source, and 10 ask for credit when the user can detect an error. On patient-facing devices, 14 submitters ask that conversational devices be evaluated across whole multi-turn exchanges, 10 ask that under-escalation and over-escalation be reported separately, and eight ask FDA not to presume that patient-facing functions are higher risk. On the generalist versus specialist question, the leading request is to base risk on the gap between the task and the user\'s competence, not on professional title.

#### Premarket evidence (Questions 7 to 17)

Competency-based benchmarking draws 49 submitters, 45 with changes. Ten ask for sequestered, contamination-controlled test sets, 10 for evidence that benchmark scores predict clinical performance, and nine for benchmarking of the final deployed configuration and not the foundation model alone. On clinical confirmation, 14 submitters ask FDA to evaluate the human-AI team in realistic workflows, and 10 would use care without the device as the comparator. On synthetic data, the requests are disclosure of the generation method and lineage (7), separate reporting of synthetic and real results (6), and a generator independent of the device\'s model family (6). On third-party testing, submitters ask for published accreditation criteria (9), conflict-of-interest rules (9), and more than one qualified provider (9).

#### Postmarket and change control (Questions 18 to 25)

No submitter accepts the trade of premarket evidence for postmarket monitoring without conditions. Of 30, 28 accept it conditionally and two reject it. Twenty would exclude high-risk functions, 18 require that failure be detectable before harm accumulates, 16 require that harm be reversible, and 13 want monitoring triggers fixed before authorization. Postmarket monitoring is the most discussed proposal (59 submitters). Seventeen caution that shared responsibility across clinicians, institutions, and societies must not dilute the manufacturer\'s accountability, and 14 want periodic reassessment combined with prespecified event triggers. On change control, 14 submitters ask for pinned, identifiable third-party model versions, 11 would define predetermined change control plans by change envelopes and acceptance criteria instead of enumerated changes, and 11 ask for rollback capability. On foundation model master files, 13 ask FDA to specify minimum content.

### My position

I have argued since the November 2024 Digital Health Advisory Committee meeting that FDA should accept lighter premarket evidence for generative AI devices in exchange for strong postmarket evidence ([my remarks are here](https://innolitics.com/articles/yujan-fda-gen-ai-ac-meeting/)). Premarket clinical evidence is the largest source of delay for the sponsors I work with, and it is a capital expense paid before a product earns anything. Postmarket surveillance is an operating cost that a prudent manufacturer carries anyway once version 1.0 ships.

The submitters are right that the trade depends on detection latency: if the first observable sign of failure is patient harm, monitoring is a record of injury and not a control. I disagree with excluding high-risk functions as a category. Risk category is a proxy for reversibility and detection latency, and some high-risk functions score well on both. I also agree with the submitters who would define a change control plan by its acceptance criteria. A sponsor cannot predict what a foundation model vendor will change, but it can fix the tests every change must pass.

### Limitations

The sample is self-selected. Trade associations, foundation model developers, and health systems have not yet filed. The coding was done by language models under rules I set, and I have not read every comment myself. Agreement was lower on the overall direction of a comment (kappa 0.70) than on positions, so that count includes only comments where both models agreed.

### Updates

The figures are regenerated nightly. I will revise the text if the distribution changes materially, and once after the docket closes.

### About the author


