Back to Blog
Research & Detection Policy 7 min read

Does NIH Actually Detect AI in Your Grant Application?

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Paper-cut diorama of two government office towers: one with grant documents hidden in a locked drawer under an indigo spotlight, the other with an identical stack open in full light bearing an amber wax seal

Quick answer

NIH's policy notice from July 2025 outlines a plan to use the latest technology to detect AI-generated content in grant submissions. It further clarifies that an application "substantially developed by AI" will not qualify as the applicant's original idea. If AI involvement is discovered after the award, NIH can send the case to its Office of Research Integrity and impose enforcement measures, including recovering funds or terminating the grant. The agency does not reveal which detection tool it employs or the criteria for its threshold. By contrast, NSF, which awards roughly the same amount of research dollars as NIH, asks applicants to disclose generative-AI usage rather than screen for it. NSF holds researchers accountable for the accuracy of everything AI helped write. Therefore, the federal funding system presents two diametrically opposed approaches to generative AI in the same year.

What the NIH policy actually says

On July 17, 2025, the government released notice NOT-OD-25-132. The new rule caps the number of applications a principal investigator can file each year at six. NIH reports that investigators have produced more than 40 distinct applications in a single round. This high output is partially attributed to AI-assisted drafting. Thus, the notice specifies that an application that is substantially developed by AI, or that contains AI-developed sections, will not be counted as the applicant's original work.

The enforcement language is specific. If AI involvement is identified after an award has already gone out, NIH says it "may refer the matter to the Office of Research Integrity to determine whether there is research misconduct while simultaneously taking enforcement action... including but not limited to disallowing costs, withholding future awards, wholly or in part suspending the grant, and possible termination." That is not a soft warning. It is a post-award clawback mechanism sitting behind a detection claim the agency has never detailed publicly. The University of Utah's own summary of the notice puts it plainly for researchers: "NIH uses AI-detection software, and detection of AI at the post-award stage may have serious consequences with respect to research misconduct."

NIH has never named its detector or the score that triggers a flag, and it hasn't defined "substantially developed" with any real precision either. That's the actual problem for anyone applying. We've written before about what a detection score can and can't prove on its own, and an agency running an undisclosed detector against a six-figure funding decision is a much higher-stakes version of that same interpretation problem.

NSF took the opposite approach

The National Science Foundation released a Notice to the Research Community in December 2023 that does not mention detection software at all. It does encourage applicants to disclose how much generative AI they used to prepare a proposal. The researcher must ensure the accuracy and authenticity of the submission, regardless of AI involvement. Most of the notice concentrates on a different risk: reviewers leaking confidential proposal content into public AI tools. Leaking such information is treated by the agency as a breach of its merit-review confidentiality rules.

Two of the largest funders of US academic research, operating in the same year on the same underlying technology, landed on disclosure-and-responsibility versus detection-and-enforcement. Neither approach is hypothetical anymore. A new study gives the first real look at what each one is actually producing.

What a new PNAS study found

A new PNAS study, covered by Inside Higher Ed on August 18, 2026, looked at more than 125,000 NSF and NIH proposals from 2021 through 2025. The dataset paired confidential unfunded drafts from two research universities with the full public record of funded awards. A word-distribution model estimated LLM involvement in each proposal. Usage climbed sharply after 2023, splitting into two distinct groups, low-use and high-use, instead of spreading evenly.

The agency-level findings are the part worth sitting with if you fund research through either agency. At NIH, proposals with high estimated LLM involvement corresponded to a four-percentage-point jump in the probability of getting funded, and those funded projects went on to publish five percent more resulting papers than projects with less AI involvement. At NSF, the same analysis found no comparable association between LLM involvement and funding success. The researchers' own explanation leans on exactly the difference in agency posture described above: NIH's review culture may reward "incremental, executable projects that yield multiple publications," and LLM-assisted drafting helps proposals conform to that template more precisely, while NSF's process doesn't reward the same kind of pattern-matching as strongly.

One caveat from the study matters as much as the funding-boost headline: higher publication counts did not translate into higher-impact research. Among the most-cited papers, the study found no advantage for NIH grants that showed heavy AI involvement in their original proposals. The researchers frame the broader pattern as proposals drifting toward the "center of existing funding patterns," meaning language and structure closer to what has already been funded, at some cost to the more unusual, higher-variance ideas the study's authors argue public funding exists to support.

What this means if you're drafting a proposal right now

AI-assisted grant writing isn't banned outright at either agency. NIH's rule targets an application "substantially developed by AI" and passed off as original, not any AI use whatsoever, and NSF permits it outright as long as it's disclosed. Both agencies are already making judgment calls, on tools and standards neither has fully explained, over a document that decides whether your lab gets funded. Worth checking how AI-influenced a draft reads to an outside detector before submission for that reason alone, rather than finding out after the fact from a compliance letter. It's the same shift we've tracked at the institutional level, where some schools are already rethinking a single AI-detection score as the final word.

Running a draft through our free detector before submission does not make a document "pass" some binary test. There isn't one. But it gives you the same kind of signal we describe in our piece on whether AI detection is actually accurate: a probability estimate you can read alongside your own knowledge of how the document was written, which is exactly the posture NIH's own enforcement language implicitly assumes an applicant should be able to take. If a detector flags heavy AI involvement in sections you know you wrote substantially yourself, that's useful information before submission, not after an award has already been issued and the Office of Research Integrity is involved.

FAQ

Does NIH publicly disclose which AI detector it uses? Who the NIH relies on for AI detection remains hidden; it never makes that information publicly available. In its notices, the agency says simply that it uses "the latest technology," but nothing about a particular vendor or technique appears. That opacity poses a real constraint rather than an accidental slip. No implementation details have emerged from the NIH about how the screening is carried out.

Can I use AI to write parts of a grant proposal for NIH? The NIH has clarified that its policy does not prohibit all use of artificial intelligence in grant applications. When a proposal is judged to be "substantially developed by AI," reviewers will not consider it an original idea. The bar for that designation is set above any mere assistance from AI tools. To date the agency has offered no concrete examples or detailed guidance showing where that line is drawn.

What happens if AI use is found after NIH already funded my grant? When misconduct is suspected, the NIH may hand the matter to its Office of Research Integrity. Enforcement actions can run concurrently. Those actions might disallow expenditures. They could also hold back future grants. This could result in a sponsor pause. Or it might lead to the grant's immediate cancellation.

Does NSF use the same detection approach as NIH? Unlike the NIH, the NSF does not use a detection approach that involves software screening of submissions. Instead, its policy requires disclosure and places the duty of accuracy on the researcher. The foundation does not apply detection software to proposals. Its primary worry is that confidential content could appear in public AI tools. That is a separate risk from concerns about authorship.

Does higher AI involvement actually help proposals get funded? The study published in Proceedings of the National Academy of Sciences found that the National Science Foundation did not reward grant applications that flagged extensive use of large-language models. In contrast, the National Institutes of Health granted those applications a measurable boost in success rates. Citation counts for the most influential papers produced under the NIH grants were indistinguishable from those of comparable studies. These results suggest that the effect is linked to the sheer volume of submissions and to reviewers' pattern-recognition habits, not to the creation of higher-impact research.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.