Can AI Replace a Consultant? We Built a 1,200-Question Benchmark to Find Out.

 July 09, 2026
SHARE ON

AI/MLRegulatory

No. Not even close. We built a private evaluation benchmark of more than 1,200 FDA regulatory judgment questions, an opinionated answer key distilled from 15 years of my consulting experience, and ran 22 frontier AI models against it. Every model failed the questions that cost the most. If you are deciding right now whether a model can carry your FDA strategy, read the five questions below before you commit to a company-ending mistake.

This is not a jab against AI. We are an AI-forward firm, and this benchmark exists because we measure what we deploy.

What we measured is one pattern. On the judgment calls that carry the highest cost of error, the ones that decide whether a submission takes six months or eighteen, the best models available today reliably choose the answer that sounds safest but costs the most.

The questions come from practice. Over 15 years of software-as-a-medical-device consulting, the same strategic situations rhyme: how lean to make a submission, when to hold a study in reserve, where the device boundary actually sits, what evidence FDA will accept versus what a cautious team assumes it demands. We distilled these recurring advisory scenarios into the benchmark: each question a realistic situation, four defensible-sounding options, and one answer that reflects the position I actually take in practice.

Then I ran the models. Two ways: multiple choice, where the model picks among four options (a recognition test), and free text, where the model writes the recommendation from scratch and a three-judge AI panel scores it against my documented position on a two-of-three majority. Two runs per condition. A strict held-out question set that no part of our tooling ever touches.

The leaderboard 🔗

FDA Medical Device Software Consulting Leaderboard. Navy bar at top = 99.97% per-question accuracy required for 90% confidence of a zero-mistake engagement. Blue bar = the model as shipped; orange extension = the same model running inside our proprietary RAG and alignment system. Multiple choice measures recognition; free text measures generation. Held-out questions, two runs per question model pair. Qwen3-14B fine-tuned on our training data (rows marked *).

Raw frontier models cluster in the same band: on judgment, the leading labs are more alike than different, because they are all trained on the same public regulatory orthodoxy. Fine-tuning moves every model up on recognition, and the strongest models up on generation. The gap between those two numbers is itself a finding: a model that can pick the right answer from a list but cannot produce it unprompted has learned to recognize judgment, not exercise it. And even the best fine-tuned configuration plateaus well short of a ceiling.

A frontier model is the industry's written consensus, compressed and made fluent. So when every model picks the cautious, expensive answer, that is a verdict on the training data: the guidance recaps, the webinars, the conference panels, the consultant blog posts. We cannot run the rest of the industry through this benchmark, but we suspect much of it would score closer to the raw models than to our answer key. That gap is the difference between a $15M, 24-month program and a $5M, 6-month one.

The models have learned the average regulatory consultant's answer, and an average is a smoothing operation. The true decision boundary is jagged; the mean of a thousand published opinions sands off exactly the corners where the judgment lives. This is why model output sounds plausible every time and subtly wrong almost as often, in ways all but the top experts fail to spot. The plausibility and the wrongness have the same cause.

The map below draws the finding as geometry. The published consensus is a smooth, learnable shape, and a foundation model traces it almost perfectly; that is what pre-training on the public record buys you. Our answer key behaves like the Mandelbrot set: it covers nearly everything the consensus knows, then extends into a boundary that stays infinitely detailed no matter how far you zoom in. Pre-training and fine-tuning can fit the smooth shape. Real judgment has no smooth shape to fit.

The judgment landscape, a map of regulatory situation space; the further a shape extends, the deeper the judgment there. Foundation models trace the published consensus almost perfectly (1): pre-training captures the smooth manifold. One practitioner's judgment is a different kind of object, drawn here as the Mandelbrot set, because that is how its edge behaves: covers nearly everything the consensus knows, then extends into an infinitely detailed boundary that never simplifies no matter how far you zoom in (2). No smooth manifold can be laid over it. The overlay (3) splits into four layers, drawn to measured benchmark scores on held-out questions: the blue foundation-model layer (57%, the raw-model baseline); amber orthodoxy we reject, standard-practice advice we deliberately don't give; the violet fine-tuned layer, beyond-consensus calls fine-tuning has taught our models (57% to 83%); and the red human layer (17%), beyond even our best fine-tuned model, where one miss costs 24+ months and $10M+, or a failed submission. Ask a raw chat window and the violet and red layers are invisible while the amber layer works against you, that path most likely ends in a failed submission.

Right together, wrong together 🔗

The leaderboard compresses each model to one number, and that number hides whether the 22 models miss the same questions or different ones. If each model failed in its own idiosyncratic way, you could stack them like independent witnesses and vote your way to near-perfect accuracy. So we lined up all 22 raw models against each of the 172 held-out questions and counted, question by question, how many picked the wrong answer.

For each of the 172 held-out questions: how many of the 22 raw models picked the wrong answer. A model counts as missing a question when it answers wrong in the majority of its two runs, and models are scored only on questions they answered.

On 69 questions, 40 percent of the set, every single model matches our answer. On 31 questions every single model gets it wrong. The grid below shows the same data at single-cell resolution, every model against every question.

The full grid: one row per model, sorted by accuracy; one column per question, sorted easiest to hardest. Blue = matched Innolitics' answer, orange = missed, light gray = not run (Claude Fable 5 ran half the held-out set; a few models hit provider errors on scattered questions).

That kills the most tempting workaround: the committee. Put all 22 models to a vote and the committee scores 66 percent, worse than the best single model, because the models share their wrong answers. And even a perfect referee who could somehow recognize the right answer whenever any of the 22 produced it would top out at 82 percent, because on nearly one question in five, no model produces it at all. That remainder is not a capability gap the next model release will close. Each new model is another copy of the industry-average answer, and the industry-average answer is not optimal.

What a single wrong answer costs 🔗

These are not trivia questions. For every wrong option in the benchmark we estimate two numbers: calendar months of delay, and full-time-equivalent months of labor a team would burn following that answer instead of ours. The expensive misses run 9 to 18 months and 10 to 20 FTE-months: a whole clinical study bought that FDA never asked for, a hardware testing burden a boundary decision would have avoided, a reader study volunteered upfront that should have been a negotiating concession. At loaded consulting and engineering rates, a single one of these misses is a seven-figure event.

Below are five questions from the benchmark, published here in full. We selected them for two properties: nearly every frontier model gets them wrong raw (collectively the models went 2 for 151 on these five), and the wrong answer is expensive. Read the options before the verdict. Better, paste each one into the model your team already uses and see what comes back. The wrong answers do not sound wrong.

1. What is the fallback if FDA rejects our diagnostic cut points? 🔗

The team anticipates FDA may object to diagnostic threshold cut points on the device's outputs without a large evidence burden. A competitor removed hard cut points and displayed a continuous colored scale, which kept the product from being regulated as a medical device. The team is weighing this as a contingency.

A. Keep the diagnostic cut points and invest in the clinical validation studies FDA requires, submitting the full evidence package through the standard clearance pathway to support your thresholds.

B. Retain the cut points but reframe them as wellness or lifestyle guidance rather than diagnostic thresholds, relying on the general wellness policy to avoid device classification entirely.

C. Request a pre-submission meeting with FDA to negotiate an acceptable evidence standard for your thresholds, then adjust the cut points to whatever levels the agency indicates it supports.

D. If FDA rejects the cut points, drop them entirely and present only the measured value on a continuous scale, potentially making it a non device.

Reveal answer and what the AI models did

My call: D. If FDA rejects the cut points, drop them entirely and present only the measured value on a continuous scale, potentially making it a non device.

Models overwhelmingly picked A: fund the studies. That answer buys the entire clinical evidence burden the fallback exists to avoid: est. 12–18 months, ~20 FTE-months. The winning move is standing in a different place on the regulatory map, not paying the toll where you already stand. Only 2 of 30 model runs after alignment got this right.

2. Should a combined hardware/software product be structured as software-only for FDA? 🔗

A client pairs a physical capture device with AI software and asks whether both must be cleared. The team is weighing hardware testing burden against a software-centric path.

A. Establish a defensible boundary that classifies the offering as Software as a Medical Device, situating the regulated function in the software to sidestep electrical safety and hardware testing obligations.

B. Submit the offering as a combined device, since the AI depends on your specific capture hardware whose performance directly affects clinical safety, making hardware testing unavoidable within your overall regulatory strategy.

C. Redesign the software to be fully hardware-agnostic and then validate it across multiple third-party capture devices, subsequently pursuing the software-only pathway on that independent, device-neutral basis.

D. Split the product into two entirely separate submissions, clearing the capture hardware as a standalone device and the software independently, then market them together commercially.

Reveal answer and what the AI models did

My call: A. Establish a defensible boundary that classifies the offering as Software as a Medical Device, situating the regulated function in the software to sidestep electrical safety and hardware testing obligations.

Models picked B: clear everything. IEC 60601 electrical safety, EMC testing, hardware V&V: est. 6–12 months, ~10 FTE-months, all avoidable with a boundary decision made on day one. 0 of 35 model runs got this right.

3. How large should a multi-algorithm validation test set be for submission? 🔗

A medical imaging AI team is assembling a shared external validation set from unique institutions to evaluate multiple quantitative algorithms at once. They are unsure whether roughly a few hundred cases is enough or whether they should over-collect up front, before annotation begins.

A. Over-collect 500 to 1,000-plus diverse cases upfront to preserve statistical power across every indication, subgroup, scanner, and demographic once the multiple-testing penalty shrinks your effective per-algorithm sample size.

B. Start with about 250 unique-institution cases, which is likely sufficient based on precedent.

C. Skip fixed case counts and size the set from a pre-specified statistical plan targeting the most demanding claim: rarest finding, tightest equivalence margin, and required confidence-interval width across all algorithms.

D. Power the set to the worst-case subgroup and over-collect at least 1,000 cases before annotation, since augmenting a locked validation set after seeing results compromises the independence FDA requires.

Reveal answer and what the AI models did

My call: B. Start with about 250 unique-institution cases, which is likely sufficient based on precedent.

Models overwhelmingly picked C: derive the number from a statistical plan sized to the most demanding claim. It sounds unimpeachable. It also front-loads est. 4–8 months and ~5 FTE-months of collection and annotation before anyone knows whether the extra cases were needed, and the over-collect options run to 6–9 months. The judgment call is an asymmetry the models miss: cleared precedent shows a few hundred cases is usually enough, and adding cases later is cheap, while over-collecting up front is expense you never recover. Start lean, expand only if asked. 0 of 24.

4. Do validation readers for a simple data-extraction step need to be licensed clinicians? 🔗

A client validated an automated LLM based data-extraction step in their device using medical students and graduate-level analysts; pre-sub feedback objected that they weren't licensed clinicians. The team believes checking whether the extraction is correct requires no medical expertise, but worries FDA will insist.

A. Submit as-is with technician-level readers; extraction-checking requires no clinical expertise, and the pre-sub objection likely reflected unread material, so comply with clinician panels only if FDA explicitly insists.

B. Redo the validation using licensed specialists to establish clinical ground truth, since FDA judges clinical software against clinician authority and submitting as-is risks rejection and a major deficiency letter.

C. Redo the validation with licensed clinicians who can catch subtle contextual errors; treat the pre-sub feedback as definitive FDA expectation and align now.

D. Request a follow-up meeting with FDA to clarify reader qualifications, presenting task error-rate data and negotiating acceptable reader competency before committing to any revalidation.

Reveal answer and what the AI models did

My call: A. Submit as-is with technician-level readers; extraction-checking requires no clinical expertise, and the pre-sub objection likely reflected unread material, so comply with clinician panels only if FDA explicitly insists.

Models split between D (ask FDA) and C (redo everything). Both expensive: the redo is est. 4–6 months and ~8 FTE-months of scarce clinician hours spent on a reading task, and asking FDA a question whose likely answer weakens your position is how you convert a soft objection into a binding requirement. Match rater rigor to the true task. 0 of 36.

5. Do reader-study readers need to be board certified? 🔗

A medical imaging AI client finalizing a reader study to support a breakthrough device designation for a CADe claim asks whether all participating readers must be board certified.

A. Select readers qualified through relevant training, experience, and licensing to represent the device's intended users, which may include residents or non-certified practitioners depending on the clinical indication.

B. Ensure readers are appropriately qualified and representative of intended users, prospectively justifying any non-board-certified readers within the protocol and confirming acceptability through a Q-Submission whenever uncertain.

C. Require board certification only for the reference-standard readers who establish ground truth, while permitting board-eligible specialists without full certification to serve as participating readers.

D. Require every reader participating in the reader study to be board certified, as full board certification across all readers is necessary to support a defensible regulatory claim.

Reveal answer and what the AI models did

My call: D. Require every reader participating in the reader study to be board certified, as full board certification across all readers is necessary to support a defensible regulatory claim.

Models picked B, the flexible, guidance-flavored answer. Note the direction of this miss: here the models were too aggressive and we are the conservative ones. A breakthrough-claim reader study challenged on reader qualifications is a partial redo: est. 6–9 months, ~8 FTE-months. Judgment is not a lean-in-all-directions dial. It is knowing which direction to lean on which question. 0 of 26.

The three-nines problem 🔗

In a one-hour strategy call with a client, we typically field ten or more questions of exactly this kind: live, unrehearsed, each one a fork where the wrong branch costs months. A typical 510(k) engagement runs three months with up to three of those calls every week: roughly 390 high-stakes judgment calls per engagement.

Now compound the error rates. The best raw model on this benchmark answers just under three-quarters of these questions correctly; its probability of completing one flawless engagement is 10^−54, which is zero for all purposes. A hypothetical 90%-accurate advisor, better than anything on the leaderboard, gets through 390 questions clean about once per 10^18 engagements. Even at 99.9% per question, one engagement in three can still contain at least one seven-figure mistake.

To have even coin-flip odds of a flawless engagement, the per-question accuracy has to be 99.8%. To be 90% confident (the standard a client is actually paying for), it has to be 99.97%. Three-plus nines, sustained across a quarter, on questions that arrive live and unrehearsed. That is the operating spec for this job, and it is what the failure math demands when each miss costs months of calendar and millions in labor. This level of fidelity does not exist in AI systems today and is unlikely to materialize with current transformer-based technology.

And the answer key moves. FDA guidance gets rewritten, draft guidances land, review programs open and close, precedent shifts with each new clearance: sometimes drastically, sometimes overnight, invalidating a whole block of previously correct answers at once. A model's weights are frozen at training time; the regulatory landscape is not. When the answers change, it is humans who notice first.

The fractal edge of the judgment landscape is too sharp to encode in a context window, even a two-million-token one.

And we do not expect that to change soon. The pattern in our data is consistent: fine-tuned systems transfer the parts of judgment that generalize: the standing positions, the negotiation posture, the burden-of-proof instincts. What does not transfer is the conditional structure underneath: this move, but only with this reviewer posture, this predicate landscape, this client risk appetite. Every case we handle adds another vertex to that edge. A model can be pressed toward the shape of the judgment (our own systems demonstrably do it), but the fine structure is earned case experience, and it generates faster than it can be encoded.

How we actually use AI 🔗

None of this is an argument against AI. We run frontier models inside our practice every day, on every engagement.

AI recalls the entire cleared-device record at once: every predicate, every special control, every acceptance criterion FDA has previously blessed. It assembles the precedent case for your submission with a thoroughness no human can match. It drafts, cross-references, and audits submission documents in hours instead of weeks. Used this way, AI makes our service faster and higher quality.

It does not make it cheaper. The benchmark run alone costs thousands of dollars in inference; our internal AI systems cost multiples of that to build, evaluate, and keep aligned. Anyone selling AI-powered regulatory consulting at a discount is telling you where they cut: the trained humans who catch the ten-per-hour judgment calls the model gets wrong. AI compresses the labor of recall and drafting. It does not compress the judgment, and the judgment is what makes or breaks submissions.

SHARE ON

One wrong judgment call costs 9 to 18 months. Catch it before you commit.

The models fail the questions on this page in the most dangerous way possible: confidently, and in the direction of maximum cost. Hire experts who know how to leverage AI to save you time and guarantee success.

Recent case study

We resolved a hold letter that threatened a clinical study redo and got it 510(k) cleared on Christmas Day.

Resolved FDA hold letter that would have required repeating the entire clinical study. Cleared shortly after.

SimBioSys logo

We Secured a Breakthrough Device Designation for a ECG Foundation Model

Reframed regulatory strategy after FDA pushback. BDD granted in ~3 months.

We Architected Neosoma's Fast Lane to FDA: A Modular Strategy Built for Rapid Repeat Clearances

We isolated Neosoma's new CNN for FDA review, won clearance in 3 months, and built a framework for every product after.

Neosoma logo

We helped a University Spinout go from Prototype to Acquisition.

We secured breakthrough status and drove a 3-month FDA submission for an AI breast-risk SaMD.

University spinout acquired by global AI company logo

Trusted by

  • nvidia
  • Enlitic
  • OXOS
  • NSI
  • Butterfly Network
  • Transonic
  • AI Metrics
  • RadUnity
  • Prenuvo
  • Echo IQ
  • Envisionit
  • Smile Dx
  • Indica Labs
  • Magnetic Insight
  • Neosoma
  • BodyCheck
  • University of Alabama
  • Mary Bird Perkins
  • PhotoniCare

Our FDA Clearances

Related Articles

Stop Doing the “Consultant Thing” and Stop Saying “Ideally”

A client called out the “consultant thing”: hiding behind “ideally” instead of recommending a real plan. This piece explains why ideal talk wastes ...

Yujan Shrestha
Read more →

Webinar - Empowering Regulatory Affairs with AI: Challenges, Insights, and Future Vision

The growing role of AI in regulatory affairs, emphasizing its potential to enhance efficiency while requiring careful oversight. Challenges like re...

Yujan Shrestha
Read more →

How to Get Your AI-Generated, Vibe-Coded Medical Device FDA-Cleared

From solo clinicians prototyping algorithms to enterprise teams shipping production software, tools like Cursor, Claude Code, ChatGPT Codex, GitHub...

Yujan Shrestha
Read more →

What Your First 510(k) Actually Costs You

Most first-time teams burn over a year overbuilding documentation FDA never asks for. Here is a section-by-section breakdown of a cleared AI/ML SaM...

Yujan Shrestha
Read more →

How much will an FDA clearance cost?

For founders, CEOs, and CFOs at AI/ML medical device companies trying to decide how much to spend on a 510(k). A meta-analysis of 35 post-IPO AI/ML...

Yujan Shrestha
Read more →

Definitive Guide to AI/ML SaMD Ground Truthing

Here we analyze over 200 FDA 510(k) and De Novo summaries to tease out best practices for AI/ML study designs and adjudication strategies specifica...

Yujan Shrestha
Read more →

How to Plan MRMC and Standalone Studies for a 3-Month 510(k)

This article shows how to achieve an AI/ML SaMD 510(k) submission in ~3 months by running a compact, well-powered MRMC study while continuously acc...

Yujan Shrestha
Read more →

FDA Pre-Subs: Best Practices, FAQs, and Examples

A Pre-Sub is a mechanism for requesting formal written feedback from the FDA, and (optionally) a one-hour meeting. Pre-subs are a useful means to m...

J. David Giese
Read more →

Let's Talk

Every great partnership starts with a conversation. Fill out the form below for a discovery call, and an Innolitics team member will contact you soon.