Back to Blog
Detection Methods & Evidence 6 min read

What Congress's Floor Speeches Reveal About How AI Detectors Spot Machine Writing

By Dusan Boljevic · AI/ML Engineer at TheChecker.AI

Torn ceremonial parchment scroll with paper-cut AI-detection marker words scattering into abstract indigo and amber ink-wash marks beside a gavel, ink-on-paper diorama style

Quick answer

Congress has a ChatGPT problem. It hides in one word: unwavering.

A 2026 study tracked 16 words and phrases across 135 million words of official U.S. Congressional speech, back to 2014. Testament. Underscore. Unwavering. The construction "not just X, but Y." After ChatGPT launched in late 2022, use of these markers spiked. Hard. The spike hit the House of Representatives the hardest. It hit hardest of all in one-minute ceremonial speeches honoring constituents. Not in real legislative debate.

Ilyass Mofaddel did the tracking. He's a University of Toronto student. He didn't touch an AI detector to find any of this. He counted words by hand. That is the same basic principle every AI detector leans on, dressed up in statistics instead of a checklist. This dataset shows exactly where that principle holds and where it falls apart.

What the study actually measured

Mofaddel published his paper in The iJournal, Vol. 11, No. 2, Winter 2026. He pulled Congressional Record PDFs straight from the Library of Congress API. He sampled 150 days a year, per chamber, 2014 through 2025. Roughly 5,000 documents. 135.5 million words. He skipped black-box AI detectors on purpose. Instead he built his own transparent measure: track 16 words and phrases, compare each year's frequency to a 2018-2022 pre-ChatGPT baseline, calculate a z-score for how far each year strayed from that norm.

His word list reads like a greatest-hits collection of AI writing tells. Testament. Pivotal. Resonate. Underscore. Elevate. Foster. Embark. Landscape. Transformative. Crucial. Tapestry. Exemplary. Delve. Unwavering. Two sentence patterns made the list too: "not just X, but Y," and "I rise to speak."

The numbers: "unwavering" jumped from a baseline z-score of -0.33 to +9.65 after 2022. "Testament" jumped from -0.62 to +6.23. "Not just X, but Y" rose from 0.35 to 5.80. Every marker combined peaked in 2024 at 12.3, then eased to 6.6 in 2025. Still far above anything seen before 2022.

The House-Senate split is the interesting part

The increase wasn't spread evenly. It concentrated almost entirely in the House. The Senate barely moved. Mofaddel ties this to how the two chambers actually work. House members represent small, local districts. They use one-minute floor speeches to honor constituents all the time. A retiring teacher. A championship football team. A fiftieth wedding anniversary. Senators represent whole states. They rarely bother with speeches like this.

Mofaddel read the 20 documents with the highest marker density by hand. He found two patterns. He calls the first "home style," a term borrowed from political scientist Richard Fenno's 1978 framework. It showed up in exactly these ceremonial tributes. One sample line: "Her many talents are a testament to her unwavering dedication, tireless commitment, and incredible passion." Three nouns. Three adjectives. You could scramble the order and nothing would change. No specific detail about the actual person being honored anywhere in it.

The second pattern turned up in floor speeches defending bipartisan legislation. Technically accurate. Never once naming a concrete mechanism of the policy.

Neither pattern showed up much in emotionally charged, partisan speeches. Mofaddel's read: politicians fighting over contested issues write with rough edges. First-person specifics. The occasional grammatical mess. What he calls "bad manners," the kind that signals authenticity to voters. Machine-generated text does the opposite. Reinforcement learning trains it to be helpful, harmless, and inoffensive, and that training smooths out every rough edge. It's built for low-stakes writing. It's terrible at sounding like someone who actually has something to lose.

This is a detection method too, just a manual one

TheChecker.AI and other AI detectors don't literally count occurrences of "delve" and call it a day. Production systems weigh probability distributions across entire token sequences, not a fixed word list. But the underlying logic sits in the same family. A 2024 analysis of 14 million PubMed abstracts found the identical thing happening in scientific writing: words like "delve," "underscore," "showcase," and "crucial" showed sudden, extreme jumps in frequency right after ChatGPT's release, while words that had shown no unusual movement in any prior year, 2013 through 2019, stayed flat. That paper and Mofaddel's Congress study never overlap, different dataset, different author, different subject entirely, yet they land on overlapping vocabulary anyway. Not a coincidence. It's what happens when millions of people start drafting with the same small set of models, trained the exact same way. Vocabulary drift is real. It's measurable. It's also blunt, and Mofaddel found a legitimate confound almost immediately: "foster" spiked hard after 2022, but a chunk of that traced back to ordinary legislative debate about foster care policy, nothing to do with AI at all. He kept the word in his index anyway, with a caveat, because that's the honest way to report a messy signal. Note what you can't rule out. Don't pretend the count is clean. Single-word or single-metric claims deserve the same skepticism. Credible detectors combine dozens of signals: perplexity, burstiness, structural patterns, and more. A wordlist alone will flag honest writers who happen to like the word "crucial," and it will always miss AI text that's been lightly edited to swap out the obvious tells.

What this means if you're the one being flagged

A few of Mofaddel's findings apply directly if your own writing has ever been flagged as AI-generated. Safe, formulaic writing with no personal detail scores worse across almost every AI-detection method, human or machine, not because it's dishonest but because it looks like the smoothed-out output RLHF training tends to produce. There's a fix, an old one: specific, first-person, occasionally rough-edged detail, the same thing that made politicians' home-style tributes read as more authentic decades before AI detection existed as a category, and the same advice educators give students trying to dodge a false flag today. Specificity reads as proof a real person wrote the thing. Run your own draft through TheChecker.AI's detector and compare what it flags against what you already know about how you wrote it.

The researcher never touched an AI detector to find any of this, and the study doesn't prove Congress is writing laws with ChatGPT either. He built a transparent, reproducible word-frequency method instead, because commercial AI detectors, in his own words, are often "opaque and lack the interpretability needed for open research." Track specific markers, compare to a pre-ChatGPT baseline, calculate the deviation. The clearest concentration of markers showed up in ceremonial, non-legislative speech and in defending bipartisan, low-controversy bills, not in contested policy fights, so the method measures a correlation in language patterns, not a confession about who actually typed what into which tool. As for why the markers eased in 2025 after peaking in 2024, and whether "delve" is a reliable tell on its own: no, "delve" is just one data point among many, and the 2025 dip has a few unconfirmed explanations. Staff and members may have drifted away from early tools that produced obviously flagged language, newer models may write with less telltale vocabulary now, and 2025's rougher political climate may have simply pushed more speech toward emotionally direct, human-sounding rhetoric. Mofaddel calls for follow-up work to see if any of it holds.

FAQ

Is "delve" actually a reliable AI tell on its own? On its own, no. It's one data point among many. See our guide on what real AI-generated text examples look like for the fuller pattern, and our breakdown of how AI text detection actually works for why detectors weigh dozens of signals together rather than any single word.

Dusan Boljevic

AI/ML Engineer at TheChecker.AI

Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.

Interested in using TheChecker.AI?

Try it free