{/* Schema recommendation: BlogPosting + FAQPage + ItemList.
- BlogPosting: author Nalin Vahil, datePublished and dateModified 2026-08-17. Emitted by the post template from frontmatter.
- FAQPage: emitted from the frontmatter faq block. Do not duplicate the FAQ in the body; the template renders it below the article.
- ItemList candidates: the four-agency comparison table in "How do the four agency innovation bars compare?" and the routing table in "Which agency's innovation bar does your idea clear?". Internal linking: first SBIR mention links to /insights/sbir-guide-for-startups; companion pieces linked inline (/post/is-your-innovation-r-and-d-or-product-development, /post/nsf-prior-art-parity-go-no-go-test, /post/nsf-decline-pattern-ai-startups, /post/arpa-h-vs-nih-decision-framework, /post/arpa-e-disruptive-innovation-bar); CTA links to /roadmap-intake. */}
"Is my idea innovative enough for SBIR?" is the question almost every first-time applicant asks. It is also the wrong question, because there is no single SBIR innovation bar.
Here is the short answer. NIH scores scientific novelty: a new mechanism, target, or method. NSF requires unproven technical risk and declines known techniques applied to new domains. ARPA-H requires a 10x, non-incremental claim. AFWERX accepts integration of mature technology if it closes a named Air Force capability gap with a defensible dual-use market.
The same idea can fail one bar and clear another. Misjudging which bar you face costs you a full application cycle: 40 to 80 hours of founder time (per Cada's engagement data), plus the 6 to 9 months most federal review cycles take to return a decline that was predictable on day one.
Cada runs a documented novelty test before drafting a single sentence of any proposal, and the test is different for each agency. We have written hundreds of proposals across 30+ agencies. This piece publishes all four tests, plus the routing logic we use to decide which agency an idea should target.
What does "innovative enough for SBIR" actually mean?
The SBIR innovation criterion is an agency-specific scored review axis, not a general quality judgment. Each agency defines innovation in its own rubric language, reviewers score against that definition, and decline letters repeat it in near-formulaic wording. "Innovative enough" therefore always means "innovative in the way this specific agency tests for."
Before asking whether your idea qualifies, you need to know what each agency's test actually measures.
This is a different question from what type of R&D you are doing. If you want to classify your project first, start with our innovation classification framework. This piece assumes you have an idea and answers the routing question: which agency's bar does it clear?
How do the four agency innovation bars compare?
The four bars, side by side:
| Agency | What "innovative" means | What fails | Typical Phase I | Who scores it |
|---|---|---|---|---|
| NIH | Scientific or mechanistic novelty: new chemistry, mechanism, biology, algorithm with a theoretical contribution, or instrumentation. A defended integration can pass. | Assembly of existing components with no named novel element and no defended integration case | Up to $323K under the standard guideline (some institutes fund $400K+), 6-12 months | Study section panel, scored 1-9 (1 is best) |
| NSF | Unproven, high-risk technical innovation. The core scientific question must be unsolved, not merely untested. | A known technique (LLM, standard ML, established CV pipeline) pointed at a new domain where feasibility is already published | Up to $305K, 6-18 months | Panel + program director, Intellectual Merit axis |
| ARPA-H | A non-incremental mechanism: 10x improvement, not 10%, on a health outcome. Something conventional approaches cannot reach. | A better implementation of an existing method, or a product description with no new funded R&D | Milestone-based awards, sized per solicitation (typically well above SBIR checks) | Program manager, against the Heilmeier Catechism |
| AFWERX | A named operational capability that closes a named Air Force gap, plus a defensible dual-use commercial case. Integration of mature technology is acceptable. | A reskin of a capability an existing program of record already delivers, or a missing dual-use case | $75K, ~3 months (Phase I) | Capability-focused evaluators, not a science panel |
Read the table and one pattern stands out: the four bars differ in how much new science they demand. NSF is the hardest test of new science. NIH is close behind but allows a defended integration. ARPA-H wants a new mechanism but frames it as capability. AFWERX explicitly does not require new science at all.
Now the four tests in detail. Each one is the actual pre-draft gate we run internally, published in full.
What counts as innovative at NIH?
NIH's Innovation criterion asks whether the application challenges existing paradigms, applies a novel theoretical concept, or uses a novel approach, methodology, or instrument. Study sections score it 1-9, where 1 is best, and they detect engineering-integration framing fast.
Our NIH pre-draft test is four questions:
- Does the proposed work contain at least one scientifically novel element? Qualifying categories: new chemistry, molecule, or construct; new mechanism, target, or pathway hypothesis; new biology or model system; new algorithm with a defensible theoretical contribution; new instrumentation or measurement modality; or a novel application of an established method where the method's behavior is genuinely uncharacterized.
- If no: what 2-3 specific technical failure modes does your integration resolve that prior approaches have not? Named failures, not vague gaps.
- Why has nobody done this integration before? Acceptable answers: a missing component, missing model system, missing measurement, or a regulatory barrier. Each named.
- What technical risk does first-time integration introduce? This risk must appear in your Specific Aims as an acknowledged risk with a mitigation.
Question 1 alone can pass the gate. If the answer is no, questions 2 through 4 must all pass together. The dominant rejection pattern we see in external reviews of SBIR drafts is "no scientifically novel content, only assembly," so an undefended integration is the thing to avoid.
An illustrative example. A startup with a new hydrogel chemistry that releases a drug in response to inflammation markers passes question 1 outright: the chemistry is the novel element. A startup assembling video visits, scheduling, and a symptom checker into a care-coordination platform has no novel element, so it needs three named failure modes, a named reason nobody has integrated these before, and a named first-time risk. That is a much harder case to defend, and most versions of it should not target NIH.
What counts as innovative at NSF?
NSF is the strictest bar, and the decline language is remarkably consistent. In the NSF decline letters Cada has reviewed, every hard decline lands on one criterion in near-verbatim words: the pitch "does not sufficiently articulate the development of a new, high-risk, technological innovation." None of those letters cites prose, market sizing, or team.
Our NSF pre-draft test:
- Where does the claimed innovation physically live? A new scientific principle, mechanism, material, or biology is low decline-risk. A known computational or engineering technique pointed at a new domain is high decline-risk.
- The AI and software test. Is the core method an LLM, transformer, retrieval pipeline, or standard ML/CV technique whose feasibility is already published, applied to a new vertical? If yes, decline risk is high by default. NSF reads this as measurement, not new science.
- The prior-art parity test. If the best published results already meet or beat your Phase I performance target on the same capability, stop. No rewording fixes this. And never cite a published benchmark as proof your own approach is feasible: the same paper cannot validate your feasibility and establish your novelty at the same time.
- Can you name one unsolved scientific question your Phase I answers? Unsolved, not untested. "How well does a known method score on our data" is a benchmarking question, and NSF declines it however clean the writing is.
Two more NSF-specific rules founders consistently miss.
Novelty-of-application is not novelty-of-method. Applying a proven technique to a domain nobody has applied it to does not create scientific novelty when the technique's feasibility is already demonstrated.
Commercial traction actively hurts your innovation argument at NSF. A shipped product with production deployments reads as low technical risk. NSF scores the R&D, not the business. We have seen decline letters that congratulate the company on its market traction and then call the proposed R&D incremental in the same paragraph. If you have a mature product, the unproven science must dominate the pitch, and the product gets one clause as a credibility anchor.
An illustrative example: a computer-vision inspection startup targets 80% defect-detection accuracy for Phase I. Its own literature review finds published zero-shot models already scoring near 90% on comparable benchmarks. That is a hard stop at NSF, no matter how good the product is. The right move is a different agency, which is exactly what the routing table below is for.
We publish the full go/no-go version of the parity test in our NSF prior-art parity guide, and our piece on the NSF decline pattern for AI startups unpacks why the software category fails so consistently.
What does ARPA-H's 10x improvement requirement mean?
ARPA-H program managers evaluate against the Heilmeier Catechism, a question set created at DARPA. The original has eight questions; ARPA-H extends it to 10 by adding health equity and misuse mitigation. The novelty question is Heilmeier question 3: what is new in your approach, and why do you think it will be successful?
The bar is non-incremental: 10x, not 10%. The improvement must be something that cannot be achieved through conventional approaches, and the PM asks whether you are offering a genuinely new mechanism or just a better implementation of existing methods.
Three rules from our ARPA-H test:
- Answer the directionality question. What does your product do today, and what would the funded R&D do that the product does not already do? If the proposal reads as a description of a finished product, it fails the non-incremental bar automatically.
- Mechanism, not adjective. State your differentiator as one mechanism a competitor cannot replicate, backed by a number, against the named current standard of care. "Better, faster, cheaper" is not a mechanism.
- Impact is health impact. Frame the payoff in lives saved, diseases prevented, or years of healthy life gained. Market size numbers signal the wrong instincts to an ARPA-H PM.
An illustrative example: a diagnostics startup claiming 15% better accuracy than the current lab standard fails the 10x test; that is optimization. The same technology reframed around a mechanism that works with no lab infrastructure at all, enabling screening for populations that currently get zero screening, is a defensible non-incremental claim. Zero to some is a 10x story; 85% to 87% accuracy is not.
Honest caveat: ARPA-H is still a young agency, and its funding patterns are less predictable than NIH's. The Heilmeier bar is documented; how consistently individual PMs apply it varies by program.
If you are weighing ARPA-H against NIH for the same technology, our ARPA-H vs NIH decision framework covers that choice in depth.
What counts as innovative at AFWERX? Capability gap plus dual use
AFWERX is the deliberate outlier. It is capability-focused, not research-focused, and integrating mature technologies into a new operational capability is explicitly acceptable. The evaluators weight operational fit, transition pathway, and commercial viability more heavily than scientific novelty.
Our AFWERX pre-draft test is four questions:
- What is the named operational capability being delivered? Name the Air Force capability gap and the specific capability that closes it. One sentence each, in operator terms, not engineering terms.
- Is the capability genuinely new to the named user? New-to-user passes. A material improvement on an existing capability passes if the delta is quantified: cost, time, accuracy, or mission readiness. The same capability an existing program of record already delivers, offered under a different vendor name, fails.
- What 2-3 specific operational failure modes does the capability resolve that the current program of record or workflow cannot? Named failures: a readiness gap, latency, a force-protection cost.
- Why is this dual-use commercially viable? A named civilian customer segment, a named pain point that overlaps with the Air Force need, and a revenue model. Generic "industry will adopt this" answers get penalized in the Commercialization Potential score.
Notice what is missing: no question about scientific novelty. The integration NSF declines is often exactly what AFWERX rewards, because AFWERX is buying capability, not science. The check size matches: Phase I is $75K over roughly 3 months, a scoping exercise rather than a research program.
An illustrative example: a startup integrating a commercial drone autonomy stack with off-the-shelf hardware for resupply in GPS-denied environments has zero new science. At NSF that is a fast decline. At AFWERX it is a strong candidate, provided the capability gap is named, the delta over current logistics is quantified, and a civilian logistics market genuinely exists.
Which agency's innovation bar does your idea clear?
The routing table we use. Find your idea profile, read across:
| Your idea profile | Primary route | Fallback | Why |
|---|---|---|---|
| New mechanism, chemistry, biology, or theoretically novel algorithm | NSF or NIH | ARPA-H if the health impact is non-incremental | You clear the hardest bars; use them, since fewer competitors can |
| Known technique, but aimed at a genuinely unsolved scientific question | NIH (defended integration) or NSF (if the unsolved question is the spine) | ARPA-H | NSF is viable only if the barrier is unsolved, not untested |
| Known technique, new domain, feasibility already published | AFWERX or other DoD SBIR | Skip NSF entirely | This is measurement, not new science; DoD rewards the integration NSF rejects |
| Mature product, operational gap, credible dual-use market | AFWERX | Other DoD programs | Traction helps here and hurts at NSF |
| Health technology with a 10x access or outcome mechanism | ARPA-H | NIH | The 10x claim is the whole game; if it is 10%, go NIH |
One routing rule worth making explicit, because it is the one our own gates enforce: an idea that fails the NSF test routes to a different agency, not to a rewrite. No amount of better writing turns benchmarking into new science.
And one honest limitation: this table is a prior, not a guarantee. Program-specific fit, topic alignment, and eligibility still have to be validated per solicitation. The table tells you where to spend your validation hours first.
One agency deliberately absent from the table: ARPA-E runs a sibling 10x bar for energy technologies. Our ARPA-E disruptive innovation piece covers how that bar is tested.
What are the most common agency misroutes?
Three patterns account for most of the wasted application cycles we see.
Misroute 1: the AI startup at NSF. A proven model class applied to a new corpus, pitched as innovation. NSF reads it as measurement and declines. Fix: route to AFWERX or DoD, where the integration and the operational user are the point.
Misroute 2: the revenue-stage company leading with traction at NSF. Deployments and customer counts signal low technical risk on the one axis NSF actually scores. Fix: if there is genuinely unproven science, make it the spine of the pitch and reduce the product to one clause. If there is not, change agencies.
Misroute 3: the research-stage idea at AFWERX. No operational user, no named capability gap, no dual-use revenue model. AFWERX scores commercialization and transition, and a science project without an operator fails those axes. Fix: route to NIH or NSF, where the science is the product.
Find out which bar your idea clears
The four tests above are the same ones we run before writing anything. If you want them run against your technology, we do a free 15-minute agency-fit screen: you answer the four questions, we tell you which agency's innovation bar your idea clears and which programs to validate first. Straight answer, no pitch, no obligation.
If the answer is "none of them yet," we will tell you that too, and what would have to be true for that to change. That costs you 15 minutes instead of a wasted 6-month application cycle.
Sources
- NIH review criteria -- the Innovation criterion language and the 1-9 study section scoring scale
- NSF America's Seed Fund -- Phase I award terms and the Intellectual Merit review axis
- DARPA Heilmeier Catechism -- the original eight questions
- ARPA-H -- the two questions ARPA-H adds to the Heilmeier Catechism (equitable access and misuse mitigation)
- AFWERX -- Open Topic Phase I award size and capability-focused evaluation
- The four pre-draft novelty tests, the routing table, the NSF decline-corpus pattern, and the 40-to-80-hour engagement figure are Cada's own methodology and internal pipeline data (hundreds of proposals across 30+ agencies)
Award figures and review criteria reflect agency guidance as of August 2026. Solicitation terms change; verify against the live solicitation before planning an application. All company examples are fictional and used for illustration only.