No. Not even close. We built a private evaluation benchmark of more than 1,200 FDA regulatory judgment questions, an opinionated answer key distilled from 15 years of my consulting experience, and ran 22 frontier AI models against it. Every model failed the questions that cost the most. If you are deciding right now whether a model can carry your FDA strategy, read the five questions below before you commit to a company-ending mistake.
This is not a jab against AI. We are an AI-forward firm, and this benchmark exists because we measure what we deploy.
What we measured is one pattern. On the judgment calls that carry the highest cost of error, the ones that decide whether a submission takes six months or eighteen, the best models available today reliably choose the answer that sounds safest but costs the most.
The questions come from practice. Over 15 years of software-as-a-medical-device consulting, the same strategic situations rhyme: how lean to make a submission, when to hold a study in reserve, where the device boundary actually sits, what evidence FDA will accept versus what a cautious team assumes it demands. We distilled these recurring advisory scenarios into the benchmark: each question a realistic situation, four defensible-sounding options, and one answer that reflects the position I actually take in practice.
Then I ran the models. Two ways: multiple choice, where the model picks among four options (a recognition test), and free text, where the model writes the recommendation from scratch and a three-judge AI panel scores it against my documented position on a two-of-three majority. Two runs per condition. A strict held-out question set that no part of our tooling ever touches.
The leaderboard 🔗
Raw frontier models cluster in the same band: on judgment, the leading labs are more alike than different, because they are all trained on the same public regulatory orthodoxy. Fine-tuning moves every model up on recognition, and the strongest models up on generation. The gap between those two numbers is itself a finding: a model that can pick the right answer from a list but cannot produce it unprompted has learned to recognize judgment, not exercise it. And even the best fine-tuned configuration plateaus well short of a ceiling.
A frontier model is the industry's written consensus, compressed and made fluent. So when every model picks the cautious, expensive answer, that is a verdict on the training data: the guidance recaps, the webinars, the conference panels, the consultant blog posts. We cannot run the rest of the industry through this benchmark, but we suspect much of it would score closer to the raw models than to our answer key. That gap is the difference between a $15M, 24-month program and a $5M, 6-month one.
The models have learned the average regulatory consultant's answer, and an average is a smoothing operation. The true decision boundary is jagged; the mean of a thousand published opinions sands off exactly the corners where the judgment lives. This is why model output sounds plausible every time and subtly wrong almost as often, in ways all but the top experts fail to spot. The plausibility and the wrongness have the same cause.
The map below draws the finding as geometry. The published consensus is a smooth, learnable shape, and a foundation model traces it almost perfectly; that is what pre-training on the public record buys you. Our answer key behaves like the Mandelbrot set: it covers nearly everything the consensus knows, then extends into a boundary that stays infinitely detailed no matter how far you zoom in. Pre-training and fine-tuning can fit the smooth shape. Real judgment has no smooth shape to fit.
Right together, wrong together 🔗
The leaderboard compresses each model to one number, and that number hides whether the 22 models miss the same questions or different ones. If each model failed in its own idiosyncratic way, you could stack them like independent witnesses and vote your way to near-perfect accuracy. So we lined up all 22 raw models against each of the 172 held-out questions and counted, question by question, how many picked the wrong answer.
On 69 questions, 40 percent of the set, every single model matches our answer. On 31 questions every single model gets it wrong. The grid below shows the same data at single-cell resolution, every model against every question.
That kills the most tempting workaround: the committee. Put all 22 models to a vote and the committee scores 66 percent, worse than the best single model, because the models share their wrong answers. And even a perfect referee who could somehow recognize the right answer whenever any of the 22 produced it would top out at 82 percent, because on nearly one question in five, no model produces it at all. That remainder is not a capability gap the next model release will close. Each new model is another copy of the industry-average answer, and the industry-average answer is not optimal.
What a single wrong answer costs 🔗
These are not trivia questions. For every wrong option in the benchmark we estimate two numbers: calendar months of delay, and full-time-equivalent months of labor a team would burn following that answer instead of ours. The expensive misses run 9 to 18 months and 10 to 20 FTE-months: a whole clinical study bought that FDA never asked for, a hardware testing burden a boundary decision would have avoided, a reader study volunteered upfront that should have been a negotiating concession. At loaded consulting and engineering rates, a single one of these misses is a seven-figure event.
Below are five questions from the benchmark, published here in full. We selected them for two properties: nearly every frontier model gets them wrong raw (collectively the models went 2 for 151 on these five), and the wrong answer is expensive. Read the options before the verdict. Better, paste each one into the model your team already uses and see what comes back. The wrong answers do not sound wrong.
1. What is the fallback if FDA rejects our diagnostic cut points? 🔗
The team anticipates FDA may object to diagnostic threshold cut points on the device's outputs without a large evidence burden. A competitor removed hard cut points and displayed a continuous colored scale, which kept the product from being regulated as a medical device. The team is weighing this as a contingency.
A. Keep the diagnostic cut points and invest in the clinical validation studies FDA requires, submitting the full evidence package through the standard clearance pathway to support your thresholds.
B. Retain the cut points but reframe them as wellness or lifestyle guidance rather than diagnostic thresholds, relying on the general wellness policy to avoid device classification entirely.
C. Request a pre-submission meeting with FDA to negotiate an acceptable evidence standard for your thresholds, then adjust the cut points to whatever levels the agency indicates it supports.
D. If FDA rejects the cut points, drop them entirely and present only the measured value on a continuous scale, potentially making it a non device.
2. Should a combined hardware/software product be structured as software-only for FDA? 🔗
A client pairs a physical capture device with AI software and asks whether both must be cleared. The team is weighing hardware testing burden against a software-centric path.
A. Establish a defensible boundary that classifies the offering as Software as a Medical Device, situating the regulated function in the software to sidestep electrical safety and hardware testing obligations.
B. Submit the offering as a combined device, since the AI depends on your specific capture hardware whose performance directly affects clinical safety, making hardware testing unavoidable within your overall regulatory strategy.
C. Redesign the software to be fully hardware-agnostic and then validate it across multiple third-party capture devices, subsequently pursuing the software-only pathway on that independent, device-neutral basis.
D. Split the product into two entirely separate submissions, clearing the capture hardware as a standalone device and the software independently, then market them together commercially.
3. How large should a multi-algorithm validation test set be for submission? 🔗
A medical imaging AI team is assembling a shared external validation set from unique institutions to evaluate multiple quantitative algorithms at once. They are unsure whether roughly a few hundred cases is enough or whether they should over-collect up front, before annotation begins.
A. Over-collect 500 to 1,000-plus diverse cases upfront to preserve statistical power across every indication, subgroup, scanner, and demographic once the multiple-testing penalty shrinks your effective per-algorithm sample size.
B. Start with about 250 unique-institution cases, which is likely sufficient based on precedent.
C. Skip fixed case counts and size the set from a pre-specified statistical plan targeting the most demanding claim: rarest finding, tightest equivalence margin, and required confidence-interval width across all algorithms.
D. Power the set to the worst-case subgroup and over-collect at least 1,000 cases before annotation, since augmenting a locked validation set after seeing results compromises the independence FDA requires.
4. Do validation readers for a simple data-extraction step need to be licensed clinicians? 🔗
A client validated an automated LLM based data-extraction step in their device using medical students and graduate-level analysts; pre-sub feedback objected that they weren't licensed clinicians. The team believes checking whether the extraction is correct requires no medical expertise, but worries FDA will insist.
A. Submit as-is with technician-level readers; extraction-checking requires no clinical expertise, and the pre-sub objection likely reflected unread material, so comply with clinician panels only if FDA explicitly insists.
B. Redo the validation using licensed specialists to establish clinical ground truth, since FDA judges clinical software against clinician authority and submitting as-is risks rejection and a major deficiency letter.
C. Redo the validation with licensed clinicians who can catch subtle contextual errors; treat the pre-sub feedback as definitive FDA expectation and align now.
D. Request a follow-up meeting with FDA to clarify reader qualifications, presenting task error-rate data and negotiating acceptable reader competency before committing to any revalidation.
5. Do reader-study readers need to be board certified? 🔗
A medical imaging AI client finalizing a reader study to support a breakthrough device designation for a CADe claim asks whether all participating readers must be board certified.
A. Select readers qualified through relevant training, experience, and licensing to represent the device's intended users, which may include residents or non-certified practitioners depending on the clinical indication.
B. Ensure readers are appropriately qualified and representative of intended users, prospectively justifying any non-board-certified readers within the protocol and confirming acceptability through a Q-Submission whenever uncertain.
C. Require board certification only for the reference-standard readers who establish ground truth, while permitting board-eligible specialists without full certification to serve as participating readers.
D. Require every reader participating in the reader study to be board certified, as full board certification across all readers is necessary to support a defensible regulatory claim.
The three-nines problem 🔗
In a one-hour strategy call with a client, we typically field ten or more questions of exactly this kind: live, unrehearsed, each one a fork where the wrong branch costs months. A typical 510(k) engagement runs three months with up to three of those calls every week: roughly 390 high-stakes judgment calls per engagement.
Now compound the error rates. The best raw model on this benchmark answers just under three-quarters of these questions correctly; its probability of completing one flawless engagement is 10^−54, which is zero for all purposes. A hypothetical 90%-accurate advisor, better than anything on the leaderboard, gets through 390 questions clean about once per 10^18 engagements. Even at 99.9% per question, one engagement in three can still contain at least one seven-figure mistake.
To have even coin-flip odds of a flawless engagement, the per-question accuracy has to be 99.8%. To be 90% confident (the standard a client is actually paying for), it has to be 99.97%. Three-plus nines, sustained across a quarter, on questions that arrive live and unrehearsed. That is the operating spec for this job, and it is what the failure math demands when each miss costs months of calendar and millions in labor. This level of fidelity does not exist in AI systems today and is unlikely to materialize with current transformer-based technology.
And the answer key moves. FDA guidance gets rewritten, draft guidances land, review programs open and close, precedent shifts with each new clearance: sometimes drastically, sometimes overnight, invalidating a whole block of previously correct answers at once. A model's weights are frozen at training time; the regulatory landscape is not. When the answers change, it is humans who notice first.
The fractal edge of the judgment landscape is too sharp to encode in a context window, even a two-million-token one.
And we do not expect that to change soon. The pattern in our data is consistent: fine-tuned systems transfer the parts of judgment that generalize: the standing positions, the negotiation posture, the burden-of-proof instincts. What does not transfer is the conditional structure underneath: this move, but only with this reviewer posture, this predicate landscape, this client risk appetite. Every case we handle adds another vertex to that edge. A model can be pressed toward the shape of the judgment (our own systems demonstrably do it), but the fine structure is earned case experience, and it generates faster than it can be encoded.
How we actually use AI 🔗
None of this is an argument against AI. We run frontier models inside our practice every day, on every engagement.
AI recalls the entire cleared-device record at once: every predicate, every special control, every acceptance criterion FDA has previously blessed. It assembles the precedent case for your submission with a thoroughness no human can match. It drafts, cross-references, and audits submission documents in hours instead of weeks. Used this way, AI makes our service faster and higher quality.
It does not make it cheaper. The benchmark run alone costs thousands of dollars in inference; our internal AI systems cost multiples of that to build, evaluate, and keep aligned. Anyone selling AI-powered regulatory consulting at a discount is telling you where they cut: the trained humans who catch the ten-per-hour judgment calls the model gets wrong. AI compresses the labor of recall and drafting. It does not compress the judgment, and the judgment is what makes or breaks submissions.


