Back to all stories
Reputational Disaster
🔴 Real Incident

Seven Months After Deloitte, EY Did It Again

16 of 27 references hallucinated. 72% of the report AI-generated. A statistic laundered from an obscure fintech blog into a McKinsey citation that never existed.

2026-05-14·7 min read·By Supervaize Team
Seven Months After Deloitte, EY Did It Again

Seven Months After Deloitte, EY Did It Again

🔴 REAL INCIDENT: EY — "Points of Attack" loyalty-fraud report retracted after AI hallucination audit (late 2025 – May 14, 2026)


What Happened

In October 2025, Deloitte Australia refunded part of a A$440,000 government contract because the 237-page assurance review it delivered quoted a federal judge who never said the words and cited academic papers that do not exist. We covered it at the time in The Compliance Review That Cited a Book That Didn't Exist. It was front-page business news across three continents. Every Big Four firm read about it.

Seven months later, EY published a report with the same defect.

The report was titled "Points of Attack: Uncovering Cyber Threats and Fraud in Loyalty Systems" — branded thought leadership released in late 2025, the kind of study a firm distributes to establish authority in a niche it intends to sell advisory hours into.

On May 14, 2026, GPTZero's investigations branch published an audit of it.

Of the report's 27 cited sources, 16 were hallucinated — fabricated outright, misattributed, or resolving to pages that had never existed. Citations were confidently credited to Forbes, McKinsey & Company, Gartner, TechCrunch, and WIRED. Following them produced missing pages and articles nobody had written.

GPTZero assessed that 72% of the study was AI-generated, and characterised the result as a "collage of misattributions, inaccurate statistics and AI-written text."

The single most instructive finding was not a fabricated source. It was a real number with a fake pedigree. One statistic appeared in the report credited to a McKinsey report. That McKinsey report does not exist. GPTZero traced the figure to an obscure fintech blog post published six months earlier.

The number had been laundered — lifted from a low-authority blog, dressed in one of the most respected names in consulting, and published under the EY brand.

EY told the Financial Times it had retracted the report and was "reviewing the circumstances that led to this article's publication."

By mid-2026, GPTZero had flagged six consulting reports with significant hallucination problems. The roster spans Deloitte Australia, a Deloitte report for Newfoundland and Labrador, EY's withdrawn study, and an AI-assisted court filing from Sullivan & Cromwell.

Both EY and Deloitte sell AI governance, AI assurance, and AI risk advisory services.


The Technical Breakdown

Citation laundering is the mechanism that matters, and it is the one nobody is controlling for. Everyone now knows a model can invent a source. Far fewer people have internalised that a model will take a real fact and attach a fabricated provenance to it, chosen for prestige rather than accuracy.

This is worse than outright fabrication in every practical respect. A wholly invented case study tends to contain something a domain expert finds odd. A real statistic credited to McKinsey contains nothing odd at all — the number is defensible, the prose is fluent, the source is exactly what you would expect to see supporting that sentence. It is undetectable by reading. It survives partner review, editorial review, and expert review, because none of those processes involve opening a citation and asking whether it exists.

And the consequence is a quiet downgrade of evidentiary quality. The client believes they are relying on McKinsey's research methodology. They are actually relying on an anonymous blog post. The claim's confidence level was silently upgraded somewhere between draft and publication, and no human made that decision.

Structured sections defeat the review process by design. Reference lists are highly patterned, which makes them trivial for a model to generate in perfect format and nearly impossible for a reviewer to assess by inspection. A partner reviewing a draft reads for argument, structure, methodology, and client fit. The bibliography gets checked for completeness, not existence. It is simultaneously the section most likely to contain fabrication and the section least likely to be scrutinised — a precise inversion of where review attention goes.

The verification bottleneck moved and nobody re-staffed it. Drafting is dramatically faster with AI assistance. Verifying 27 citations takes exactly as long as it did in 2019. A firm that captures the drafting speedup without expanding verification capacity has built a pipeline that structurally ships unverified claims — not occasionally, but as its normal operating mode. This is the same arithmetic driving the AI citation sanctions wave through the legal profession, where courts now itemise penalties at $500 per invented case.

Detection came from outside, again. GPTZero found this. Journalists at the FT amplified it. EY's own quality control — mature, expensive, thoroughly documented — did not. That is not because EY's reviewers are careless. It is because those processes were built to catch analytical weakness, methodological gaps, and client-relationship risk, all under an assumption that held for a century: a citation appearing in a draft means a human found something and read it.

The audit GPTZero ran required no privileged access. EY could have run the identical check on its own report, before publication, for a rounding error against the engagement value.


The Broader Pattern

The uncomfortable fact is not that a Big Four firm used AI. Everyone is using AI, the efficiency is real, and pretending otherwise would be theatre.

The uncomfortable fact is that the warning did not work.

Deloitte Australia was not an obscure incident. It was a public, expensive, internationally-reported demonstration of exactly this failure mode, with a named client, a refunded fee, and a clear causal explanation. It was the best possible industry-wide training signal: here is the defect, here is what it costs, here is the control you are missing.

Seven months later a peer firm shipped the same defect. Not a variant — the same one. Fabricated and misattributed citations in a published client-facing report, caught by an outsider.

That tells you something important about how organisations actually respond to incidents at other organisations. The Deloitte story was almost certainly discussed at EY. It was probably in a risk committee deck. What evidently did not happen is the thing that would have mattered: someone changing the pre-publication checklist so that every external reference had to be opened and confirmed by a named person before the report shipped. Awareness is not a control. It never has been.

There is a second-order harm worth naming. Consulting reports are inputs to other people's decisions. Governments cite them in policy. Boards allocate capital against them. Competitors quote them in strategy documents. Journalists repeat their statistics. That laundered figure did not stay inside the EY report — it entered the citation graph and began acquiring legitimacy through repetition. The retraction reaches a small fraction of the audience the original reached. Some of those numbers are now permanently in circulation, sourced to a McKinsey report that was never written.

The regulatory layer is closing in behind. FINRA's 2026 Annual Regulatory Oversight Report added a dedicated generative AI section naming hallucinations as a risk firms must test for and govern. McKinsey's own 2026 AI Trust Maturity Survey found inaccuracy to be the most-cited AI risk among people with direct responsibility for AI governance. The profession knows. The controls are behind the adoption, and the gap is now measured in public retractions.


How It Could Have Been Prevented

  • Open every external reference before publication, with a named verifier and a record. Not a formatting pass. Not a spot check. Each citation resolved, opened, and confirmed to say what the text claims. For a 27-source report this is a few hours of work — and it is the assurance being sold.
  • Escalate verification priority for prestigious attributions, don't relax it. A cite to McKinsey, Gartner, Forbes, or WIRED is exactly what a model produces when it needs authority for a plausible sentence. Treat high-status sources as high-risk until confirmed, which is the opposite of how reviewers instinctively weight them.
  • Run citation-audit and AI-detection tooling on deliverables pre-publication. GPTZero did this from the outside with no special access. Any firm can run the same analysis on its own drafts, and the cost is negligible against the engagement fee.
  • Give the reference list its own reviewer, separate from the drafter. Highest-risk section, lowest-scrutiny section. It needs a dedicated owner with a dedicated checklist, not a completeness glance from the partner reading for argument.
  • Disclose AI use in the engagement up front. Deloitte disclosed Azure OpenAI use after the errors surfaced, which reads as explanation rather than transparency. Disclosed in advance, it becomes a methodology note the client accepts — and it forces an explicit conversation about what verification the client is buying.
  • Convert peer incidents into checklist changes, with a deadline and an owner. The real failure here is organisational. When a competitor suffers a public, well-documented failure in a process you also run, the response cannot be a briefing. It has to be a diff to a control document, assigned to someone, with a date.

The Lesson

Professional services firms sell one thing underneath all the frameworks and methodologies: the assurance that qualified people checked.

A client could research loyalty-program fraud themselves. What they cannot produce is work carrying the institutional weight of a Big Four name — the implicit warranty that the firm's process, reputation, and liability exposure stand behind every claim in the document. That warranty is the entire product. Everything else is formatting.

AI does not damage that product by writing badly. It writes well. It damages it by generating content indistinguishable from verified work while being unverified, and by concentrating that content in precisely the sections review processes were built to skim.

What makes the EY case sharper than Deloitte's is that it came second. Deloitte can fairly claim it was early — the failure mode was not yet common knowledge, the controls were not yet obvious, the industry had no worked example. EY had the worked example. It had seven months, a named victim, a public post-mortem, and a refunded fee to point at.

The report still shipped with 16 of 27 sources that did not survive a check any intern could have run.

That is the part worth sitting with. The gap between knowing about a failure mode and being protected from it is an explicit, staffed, documented control — and almost nobody builds one until the retraction has their own name on it.

Take your last client deliverable. Pick five external claims at random — a statistic, a citation, a market figure, a case reference, a quote. Open each source and confirm it says what you said it says. If you cannot complete that in twenty minutes, you are shipping the same unverified assurance EY shipped, and the person who eventually checks will not work for you.


Sources

  • GPTZero Investigations — "Investigation: Hallucinations in Ernst & Young Report on Loyalty Fraud," May 14, 2026
  • Information Age (ACS) — "EY retracts cyber report littered with AI errors," 2026
  • Computing.co.uk — "EY cybersecurity report pulled after probe finds 'AI hallucinations'," 2026
  • Consulting.ca — "EY Canada takes down study after apparent AI hallucinations," 2026
  • Sherwood News — "AI hallucinations appear to be creeping into consulting reports," 2026