He Trusted AI — Then Lost EVERYTHING

The lesson of the Deloitte affair is not that artificial intelligence lies — it is that the workflows built around professional reporting were never designed to catch a machine that lies fluently, confidently, and at scale.

Key Points

  • Deloitte’s Australia welfare report contained more than 20 fabricated citations and a misquoted federal court judgment, later traced to an undisclosed generative AI tool.
  • A named academic, Chris Rudge, caught the errors; another, Lisa Burton Crawford, confirmed a book attributed to her simply does not exist.
  • Deloitte issued a corrected report and refunded part of its fee, but disputes the claim that AI alone drove the failure and stands by its substantive conclusions.
  • Similar fabricated-citation incidents have since surfaced in Canadian, European, and American government-commissioned reports, suggesting a systemic pattern rather than one bad document.
  • Vendors and risk advisers counter that hallucination is a manageable engineering problem — grounding, citation requirements, and human review — not proof AI reporting is inherently untrustworthy.

What Actually Went Wrong Inside the Report

In 2025, Deloitte delivered a report to Australia’s Department of Employment and Workplace Relations, commissioned to help the government tighten its welfare compliance and automated penalty system. The fee was reported at roughly A$440,000. When Sydney University law lecturer Chris Rudge read it closely, he found more than 20 discrete errors: references to academic works that do not exist, and — more damaging for a document meant to justify legal enforcement mechanisms — a misquoted Federal Court judgment. Academic Lisa Burton Crawford went further, publicly stating that a book the report attributed to her, on administrative justice and Centrelink, was never written by her at all.

These were not typos or sloppy formatting. They were invented sources dressed in the full apparatus of scholarly citation — author, title, plausible subject matter — footnoting claims in a document meant to guide government policy on welfare enforcement. The department confirmed the report contained “glaring errors in footnotes and sources” and demanded that consultants disclose AI use and tighten quality assurance going forward. Deloitte’s revised appendix later admitted that “a part of the report included the use of a generative AI large language model licensed by DEWR and hosted on DEWR’s Azure tenancy”. The corrected version stripped out the fictitious references and rewrote the bibliography. The government accepted a partial refund, reported at roughly A$97,000.

Why This Keeps Happening: The Mechanics of a Hallucination

A large language model does not retrieve facts the way a database does; it predicts the statistically likely next word based on patterns in its training data. Ask it for a citation supporting a claim, and if no perfect match exists in its memory, it will often generate one that looks exactly right — correct academic tone, plausible author name, a journal that sounds real — without any internal signal telling it the reference is fabricated. This is the phenomenon researchers call hallucination, and commentary on the Deloitte case explicitly described the invented court quote and non-existent references as “classic examples of an AI hallucinating”. The danger is structural: consulting deliverables reward speed, breadth of citation, and polished prose, which are precisely the conditions under which a hallucinating model is most likely to slip past a rushed human reviewer.

What the public record cannot yet establish is the precise division of responsibility between the model and the humans who used it. Deloitte has maintained that it stands by the report’s substantive findings even while correcting the citations, and it has disputed characterizations of how forthcoming it was about AI use before the errors surfaced. That is a live, genuine dispute — the fabrications are undeniable and documented, but whether the AI alone produced them, or whether human drafters copied and preserved bad references, has not been conclusively separated out in anything made public so far.

Not an Isolated Incident

The temptation is to treat the Australian report as an unfortunate one-off. The evidence does not support that reading. Months later, a Deloitte-authored Health Human Resources Plan for Newfoundland and Labrador, costing the province roughly $1.6 million, was found to contain fabricated academic citations that researchers said resembled AI-generated hallucinations. In the same province, a separate Education Accord report — a 410-page document costing $750,000 — was pulled from the government’s website after professors discovered nearly three dozen references to sources that did not exist. Reporting has since linked fabricated AI citations to policy documents tied to ENISA and the White House, and a GPTZero investigation found that 60 percent of the references in a December advisory report from EY on loyalty-program cybersecurity appeared to be hallucinated. Even a report meant to justify Australia’s under-16 social media ban was found to contain fabricated citations traced to ChatGPT. The pattern is now broad enough that no single firm or country can be treated as the anomaly.

The Counter-Case: Is This a Solvable Engineering Problem?

The strongest pushback against treating this as an indictment of AI-assisted reporting comes not from denial that hallucinations occur, but from a body of vendor and risk-advisory guidance arguing the failure mode is manageable. PwC advises firms to build “hallucination-specific controls and tests,” including grounding models in verified data and testing multiple prompt scenarios before releasing output. Other advisers recommend retrieval-augmented pipelines that tie every claim to an approved internal document, and insist that any answer the system cannot support with a retrievable source be flagged for human review before it reaches a client. The consistent refrain across this material is that generative AI “works best as a first draft, not a final authority” — a framing that treats the Deloitte failure not as proof AI cannot be trusted in professional work, but as proof that the verification layer between AI output and publication was, in this instance, absent.

That argument has real force, and it deserves to be taken seriously rather than dismissed as industry spin. But it does not contradict the core facts of what happened — it explains how it could have been prevented. The existence of a plausible fix does not erase the fact that a taxpayer-funded government report on welfare enforcement shipped with a fabricated judicial quotation and more than 20 invented sources, was caught by an outside academic rather than internal review, and had to be corrected after publication.

What It Means Going Forward

Government procurement is now visibly catching up to a problem it did not anticipate. Australia’s demand that consultants declare AI use and maintain quality assurance is an early, ad hoc response to a gap that had no standard before these incidents forced the issue. Until disclosure and audit requirements become routine parts of consulting contracts — not an afterthought triggered by embarrassment — buyers of high-stakes reports, whether government agencies or corporate boards, are the ones absorbing the risk that a footnote might not exist at all.

Sources:

youtube.com, hrreporter.com, fortune.com, incidentdatabase.ai, amjid.au, linkedin.com, reddit.com, horsesforsources.com, afr.com, koreadeep.com, ud.hk, pwc.com, techjacksolutions.com, cbc.ca