PwC's AI-Written Report Invented an Entire Product. Four Governments "Used" a Framework That Doesn't Exist.
By Dusan Boljevic · AI/ML Engineer at TheChecker.AI
Quick answer
An AI detection group's investigation, verified by the Financial Times, found that PwC Middle East published four "thought leadership" reports between 2024 and 2026 riddled with fabricated citations and invented facts. The worst one described a PwC framework called "Citizen Pulse" and claimed the governments of Denmark, Saudi Arabia, the United States, and Australia were already using it to run public services. No such framework appears to exist anywhere else, and none of the cited sources back the claim. PwC isn't alone. The same pattern hit KPMG, EY, and Deloitte (twice, including a report a Canadian province paid $1.6 million for) over the past year. None of this happened because a company used AI to draft a report. It happened because nobody checked the citations before publishing it under the firm's name.
What was actually in the reports
The report in question, "Transforming Governance," described Citizen Pulse as a "dynamic mechanism" and an "innovative approach" for governments to gather real-time citizen feedback, without ever explaining what it actually was or how it worked. It claimed Danish public institutions were already using Citizen Pulse data, that Saudi public institutions used it "to drive public service reforms," and that federal and state service portals in Australia relied on it too. Investigators checked the cited sources for each claim. Denmark's citation didn't mention Citizen Pulse. Saudi Arabia's citation was the country's own Vision 2030 page, which says nothing about a PwC product. Australia's citation was a press release that never mentions Citizen Pulse, Veterans Affairs, or the agency the report attributed the claim to. The tells in the other three reports were just as blunt, according to the Financial Times. One footnote cited a PwC survey claiming 70% of Middle East CEOs expect generative AI to reshape their business, but the linked article never mentions any survey. Another footnote's web address still carried the tracking tag "utm_source=chatgpt.com." A cited academic paper on air quality in Riyadh appears to have been invented outright: no trace of it exists in the journal it was attributed to, or from the authors it names. PwC told the FT it "takes the accuracy of our published research seriously" and was "updating a limited number of supporting citations." The Citizen Pulse report has since been pulled from PwC's website entirely.
This isn't a PwC problem. It's a pattern, and a detection score alone wouldn't have caught it.
PwC isn't the first Big Four firm caught doing this. It's the fourth in roughly a year. KPMG pulled an October report after false claims about AI deployments at UBS, the NHS, and Transport for London turned out to rest on fabricated case studies. EY retracted a study over fake footnotes. Deloitte Australia partially refunded the Australian government roughly AU$440,000 once a commissioned report was found citing academic papers that don't exist, plus a quote falsely pinned on a federal court judge. Deloitte's Canadian arm did it worse: a $1.6 million Health Human Resources Plan for Newfoundland and Labrador cited at least four research papers that don't exist, including one credited to a real nursing researcher, who told a local reporter on the record that the paper "does not exist" and that her team never had the financial data behind the analysis the report claimed she'd published. The same province had already been burned once before, by a fabricated-sources education policy paper. Two hits, one government, a few months apart. Different firms, different dollar amounts, same mechanism: a model fills a gap with a plausible-sounding citation, statistic, or case study that isn't real, and nobody downstream checks the sources against the claim before it ships under a name everyone trusts. It's tempting to read a story like this and assume an AI detector would have flagged the fake framework immediately. That's not quite right, and it's worth being precise about what a detector actually measures. An AI-writing score estimates the statistical likelihood that a passage of text was generated or heavily edited by a language model. It does not verify that a citation says what a report claims it says, and it does not confirm that a named product, government partnership, or academic paper actually exists. Reporters caught the Citizen Pulse fabrication by clicking through to the primary sources and reading them, the same manual step that should happen before any client-facing document ships. The same investigation that flagged the PwC reports as heavily AI-written also had to do that source-by-source legwork to prove the framework didn't exist; the writing-style score and the fact-check are two separate jobs, and this newsroom's own explainer on why a high detection score is not proof of anything makes the same point from the opposite direction: a score is evidence worth investigating, not a verdict on its own. What a detection pass is genuinely good for here is triage. A report that scores as heavily AI-generated or AI-edited deserves a slower, more skeptical read before it goes out, the same way a smoke alarm doesn't tell you where the fire is but tells you to go look. Investigators in these cases have repeatedly found a real correlation between high AI-writing scores and the presence of fabricated citations in the same document, which is exactly the kind of early-warning signal a comms or research-review team can act on before a report reaches a client's inbox or a government's website, rather than after a reporter finds it.
The same failure mode shows up everywhere AI drafts unchecked prose
AI-generated drafts passing review is not a quirk of corporate consulting; it has already appeared in academic papers submitted to peer review. In courtrooms, lawyers have received sanctions or fines for using AI-invented cases and quotations in their briefs. Each case shares one common element: a trusting, untested relationship with machine output. The solution is not to stop using AI tools. Instead, regard any AI-produced material as a first draft that still demands a human hand to confirm each fact against its real source before presenting it as your own.
What this means if your team publishes anything under a brand name
If your organization produces client-facing reports, research summaries, press materials, or any document meant to carry institutional credibility, the PwC and Deloitte incidents are a preview of what happens when that review step gets skipped. A practical process looks like this: run the draft through an AI detector as an early tripwire, not a verdict, treat a high score as a signal to slow down and check every citation by hand, and never publish a stat, case study, or named partnership without someone actually opening the source and confirming it says what the draft claims. That last step is the one every one of these incidents skipped, and it's the only one that actually would have caught a fabricated product being credited to four real governments. Want to know if a draft carries the same red flags that got these reports pulled? Run it through our detector and read the sentence-level breakdown instead of trusting a single pass-fail number.
FAQ
Did using AI cause these reports to contain false information, and would a detector have caught the fake "Citizen Pulse" framework before publication? The fabrication pattern here (invented citations, invented case studies, an invented product) is specific to how generative models fill a gap with something plausible-sounding when no real source exists. It's not the same failure mode as a human writer padding a report or making an honest mistake; the model produces confident, correctly formatted text with no basis in reality, and the failure that let it ship wasn't the AI use itself, it was nobody checking the output against real sources before publication. A detector wouldn't have caught it directly either. It flags the statistical likelihood that text was AI-generated, which correlates with fabricated-citation risk in these documented cases, but confirming a named product or government partnership doesn't actually exist still requires someone reading the cited sources and checking by hand. A high score is a strong reason to slow down and do that check, not a substitute for doing it.
Is this just a PwC problem, and what's the actual fix beyond "check your work"? No. Over the past year, KPMG, EY, and Deloitte have all been required to retract or compensate for reports with the same fundamental flaw, across different countries and report types, from AI strategy papers to a $1.6 million government health initiative — a sector-wide issue, not a PwC-only one. The fix is building the check into the workflow instead of trusting individual memory: every draft carrying the organization's name goes through an AI-detection pass as a fast first filter, then a named person manually verifies every citation, statistic, and named partnership against its real source before the document is approved to publish. Every report that got caught skipped that second step entirely.
Dusan Boljevic
AI/ML Engineer at TheChecker.AI
Dusan Boljevic writes at TheChecker.AI, covering how AI-text detection works and how students, writers and teams can use it responsibly.
Related posts
What Congress's Floor Speeches Reveal About How AI Detectors Spot Machine Writing
A 135-million-word study of Congress shows which vocabulary tells give away AI-written speeches, and why some lawmakers use far more of them.
Read more
Should a Single AI-Detector Score Ever Decide a Case?
A fresh CACM feature on a College of Charleston lawsuit shows what actually loses in court: not a wrong score, a missing conversation.
Read more
Can an AI Detector Catch a Fake Legal Citation Before It Gets You Sanctioned?
A law firm representing a major bank just got its brief struck for fake AI citations. Here's what a detector can and can't catch in a legal filing.
Read more