The Deloitte AI report refund: anatomy of a missing review step
In 2025, Deloitte agreed to partially refund the Australian government for a AU$440,000 report that contained a fabricated quote from a federal court judgment and references to academic papers that do not exist, produced with the help of a GPT-4o toolchain. The errors sailed past a major firm's quality process and were caught by a single welfare academic who noticed papers attributed to his colleagues that they had never written. That is the whole lesson in one story: the defense against confident AI errors is not more polish or a bigger brand, it is a reader close enough to the sources to check them. This post walks through what happened and turns it into a review step you can run before your next deliverable.
Updated

A AU$440,000 report with invented footnotes
In December 2024, Australia's Department of Employment and Workplace Relations commissioned Deloitte to run an independent assurance review of its Targeted Compliance Framework, the automated system that penalises jobseekers who miss welfare obligations. The 237-page report was published in July 2025, and it looked exactly like what a top-tier firm delivers: structured, referenced, fluent.
Inside the polish sat fabrications. A quote attributed to a federal court judgment that the judgment does not contain. Citations to academic papers that were never written, attributed to real professors at Lund University and the University of Sydney. Deloitte later confirmed that a generative AI toolchain based on Azure OpenAI GPT-4o had been used in producing the report, and agreed to repay the final instalment of the contract. A corrected version, with the AI use disclosed, replaced the original.
Who actually caught it
The errors were not found by a procurement checklist, a second AI, or the firm's own review layers. They were found by Chris Rudge, a welfare academic at the University of Sydney, who read the report closely and noticed something specific: papers attributed to his own colleagues that he knew did not exist.
That detail is worth sitting with. Every institutional safeguard between a global consultancy and a national government let the fabrications through, because at every layer the report was reviewed for how it read, not for whether its sources held. The one reader who checked was the one who knew the terrain well enough to be surprised. Verification did not fail because it is hard. It failed because nobody upstream considered it their job.
Why the polish passed and the substance did not
Nothing about this failure is unusual to AI. It is the pattern we have written about across this series: models restate weak or invented material in the same confident register as solid material, so a fabricated footnote reads exactly like a real one. Fluency is what review processes are tuned to catch, and fluency was flawless.
The economics made it worse. Checking a claim means opening the source behind it, and a 237-page report carries hundreds of claims. Without tooling that makes each check cheap, the honest cost of full verification looks prohibitive, so organisations quietly skip it and hope. The Deloitte case is simply what that hope looks like when it lands on the front page, with a refund attached.
The review step that would have caught it
The fix is not to ban AI from deliverables. The report's problem was not that AI helped write it; it was that no cheap path existed from each claim back to its evidence, so nobody walked that path before the client did.
A working review step has three parts. First, provenance by construction: every factual claim in the draft carries a link to the exact passage it came from, so checking is a click, not a research project. Second, a verification pass scoped to what matters: the claims you would be most embarrassed to retract, the numbers that drive the recommendation, anything surprising. Third, ownership: the person whose name is on the deliverable runs the pass, because the Deloitte story shows what happens when everyone assumes the layer below did it.
On that last point, the story generalises beyond consulting: whoever signs, checks. It is the same rule teams use for any AI-assisted work product.
Before your next AI-assisted deliverable
This is the posture Tatsulok is built around. Answers come only from your documents and curated collections, every claim links to its exact passage, and gaps are stated honestly instead of filled with plausible inventions. The review step that was missing in the Deloitte report is the interface: you read the claim, you open the passage, you decide. Proposals stay proposals until a human accepts them.
This week's drill, no signup needed. Take one AI-assisted document you or your team shipped recently and pick the three claims that would hurt most if they were wrong. Trace each one to its source. If all three trace cleanly, you have a review habit worth keeping. If one does not, you have found the same gap Deloitte found, at a much better price.
FAQ
- What happened with the Deloitte AI report in Australia?
- A AU$440,000 assurance review Deloitte delivered to Australia's Department of Employment and Workplace Relations in July 2025 was found to contain a fabricated federal court quote and citations to nonexistent academic papers. Deloitte confirmed a GPT-4o based toolchain was used and agreed to repay the final instalment of the contract.
- How were the errors in the Deloitte report discovered?
- Not by an internal process. Welfare academic Chris Rudge of the University of Sydney read the report closely and recognised that papers attributed to his colleagues did not exist. A domain expert checking sources caught what every institutional review layer missed.
- Did Deloitte refund the full amount of the report?
- No. Deloitte agreed to repay the final instalment of the AU$440,000 contract, a partial refund, according to the department's statement. A corrected version of the report with the AI use disclosed replaced the original.
- Does this mean AI should not be used for professional reports?
- No. The failure was not AI assistance; it was the absence of a cheap path from each claim to its evidence, so no one verified before delivery. AI-assisted work with passage-level citations and a scoped verification pass by the person who signs is a defensible workflow. Unverified fluency is not.
- How do I prevent fabricated citations in my own AI-assisted work?
- Work from grounded tools that link every claim to its exact source passage, then verify the claims that carry the most risk: the ones driving decisions and the ones you would have to retract publicly. The rule that survives this story is simple: whoever signs the deliverable runs the check.
Sources
Related guides

Hebbia alternative
Looking for a Hebbia alternative? Tatsulok gives cited, verifiable answers from your documents with team sharing, without enterprise per-seat pricing.
Read the article
Harvey AI alternative
A Harvey AI alternative for teams that want cited, verifiable answers from their documents across domains, not only legal, with access-controlled sharing.
Read the article
AI document review
In a landmark benchmark, AI reviewed NDAs at 94% accuracy in 26 seconds versus 85% and 92 minutes for experienced lawyers. What AI document review does well, where it fails, and how to run it safely.
Read the article