peakstateglobal Work with us

Three references in a A$440,000 government report did not exist

The report had been through every review such a contract carries. What that episode actually reveals about your own sign-off chain, and the seven checks that close it.

The slide had a subscriber number on it that I could have recited in my sleep. I had presented it at least a dozen times, to rooms with senior people in them, and on a quiet Tuesday in August I opened the source behind it for the first time. The real figure was three times bigger than mine. I had been underselling my own strongest case, confidently, for months.

Over the next two days I found eleven errors like that, across three of my own workshop decks. All three had been reviewed. One had already been through a full correction pass, which means the errors in it survived a review whose entire purpose was to find errors.

If you have shipped anything AI-assisted in the last year, I suspect you have a version of this feeling you have not said out loud. You reviewed it. You read the whole thing, you followed the argument, you flagged the two sentences that landed oddly, and it went out. And somewhere underneath that, you know you did not open a single source. I hadn't either. Neither has anyone who reviews our work, and nobody has ever asked any of us whether we did.

The obvious reading of my eleven is that I was careless. I sat with that one for a while, because it is not a comfortable number to own. And yet it is the least useful thing the number can tell you. Every one of the eleven had been read, by me, more than once. They all read fine. They read fine because I wrote them, and I write things that read fine. That is the actual finding: reading is not a check, and almost everything we call review is reading.

What were the eleven, exactly?

A subscriber number understated threefold, which meant the deck was underselling its own strongest case. A ninety-five per cent finding attached to the wrong denominator: it counted organisations, and I had pinned it to a set of three hundred initiatives that was a different part of the same study. An analysis credited to the wrong author, because I had cited the blog post that pointed at the analysis rather than the analysis itself, and the blog post contained no numbers at all. A benchmark quoting a news article's rounding rather than the paper's own figure, and using the wrong task set as the denominator while it was at it. A sentence in quotation marks, attributed to a named journal, that no index I can reach has any record of.

And, in a document arguing that you should not fabricate attributions, a fabricated attribution. Not my finest moment.

Did you notice what is not in that list? Nothing is a lie. Nothing is even really a mistake in the ordinary sense. Every single one is a claim that drifted a small distance from its source and kept its confident shape all the way to the slide. I have started calling this Drift, because once you name it you start seeing it everywhere, including in the work of people far more careful than me.

Why didn't reading catch them?

Because a drifted claim looks exactly like a correct one. There is nothing on the page for a reader to notice.

In August 2026 the count of court decisions involving fabricated AI-generated citations read 1,922 on the sixteenth and 1,961 on the twenty-sixth. Thirty-nine in ten days, in courts, where somebody was paid to check. I can tell you the number moved, rather than just assert it, because I captured the page both times and hashed the text.

A deck of mine quoted "past 1,500". By the twenty-sixth the real figure was 1,961, and not one character of the deck had changed. The claim rotted quietly underneath the document, and nothing in the document could tell you.

This is not new, and it is not really about AI. Zittrain, Albert and Lessig measured it back in 2014 and found more than seventy per cent of URLs cited in three legal journals, and fifty per cent of those in United States Supreme Court opinions, no longer produced the information originally cited. Not dead links. Live links, pointing at something that no longer says the thing.

I should be careful with that figure rather than let it do more work than it can. It is a study of legal citations specifically, more than a decade old, so the exact percentages do not transfer to your board pack. What transfers is the mechanism. My own view, and it is a view rather than a finding, is that it has got worse since, because far more of what we cite now is a live page rather than a printed volume.

None of this started with AI. It made every part of it faster, cheaper and harder to see, at a volume nobody reads.

One disclosure before we go on: I publish the framework this article is about, so weigh my eleven accordingly.

Deloitte refunded a report because nobody opened the references

In 2025 Deloitte Australia delivered a A$440,000 assurance review to the Department of Employment and Workplace Relations. Three of its academic references turned out not to exist. The firm issued a partial refund and republished the report with a disclosure that a generative AI system had been used in writing it.

The reporting called it an AI story. I think it is a procurement story.

The department did not buy a document, it bought an assurance. What the episode actually revealed is that nobody in the chain had been asked what they personally checked, and that when the question was finally asked, it was asked by an academic reading the report afterwards for his own reasons. The report went through the reviews such a contract normally carries. Whatever those reviews were, none of them opened the references.

Now put your own organisation in that sentence. If a figure in last quarter's board pack turns out to be invented, who answers for it? Not "which process". Which person, and what did they check?

If the honest answer is that the question has never been asked, you are the default. So was I.

"Get a human to review it" is right, and it is not the answer

I want to be careful here, because the obvious fix is not wrong.

Human review is correct. It catches things nothing else catches. And yet it is also what every organisation already does, and it is what every one of my three decks had already been through. The advice cannot be the fix, because the fix was already in place when the eleven errors went out.

The gap is not that nobody reviewed it. The gap is that "review this" does not say review it against what. A signature buys accountability. It does not buy accuracy, and pretending otherwise is how a chain of people each assume the checking happened one step earlier. We have all been a link in that chain, me very much included.

What a reviewer needs is not more diligence. It is a question narrow enough to answer.

The seven ways it goes wrong

After the eleven, I went looking for the shape of the problem rather than more examples of it. Everything I could find, my decks, the Deloitte report, the court cases, kept reducing to the same seven failures. None of them is new, and none of them is about AI. They are human failure modes older than the tool. What AI removed is the friction that kept them small: the review boards, the forums, the second pair of eyes all quietly relied on somebody reading every output, and that assumption is gone.

Here they are, in the order they bit me:

i) Nobody can tell where any of it came from. My eleven, one by one. Fluent prose stands in for provenance, and fluent is exactly what these tools produce. ii) Nothing ever argued back. Every draft agreed with me. Every review agreed with the draft. Nowhere in the chain did anything make the strongest case against a claim before it shipped. iii) If it is wrong, no name is on it. The Deloitte question, which person and what did they check, had no answer, because nobody had ever been assigned one. iv) The reasoning was never written down while it happened. Ask me a month later why a figure was in the deck and I will give you a confident story. A reason reconstructed afterwards is a story, not a record. v) There is no safety net. Nothing in my process stopped a number appearing without a source, or a quote appearing without an original. The lines existed in my head, which is the same as not existing. vi) The thing changed since it was checked. My "past 1,500" was true when I wrote it. The world moved to 1,961 and the deck had no way of knowing. vii) The reader cannot tell what was checked, or by whom. Every document I have ever sent looked equally finished, whether I had verified everything in it or nothing.

Where does your own last quarter of output sit against those seven? My guess is at least five of them are running in it right now, unwatched. That was my count, and I do this for a living.

Seven checks, one each

The framework I now run on everything that ships under my name is called SOURCED, and each letter is one check answering one of the seven:

  • Sourced. Every load-bearing claim names a source that was actually retrieved, with the verbatim sentence and where it lives. A claim you are certain about, but did not go and look at, is labelled as memory, never as fact.
  • Opposed. Before it ships, something argues back: the strongest case against each conclusion, and what survived it.
  • Underwritten. A named person answers for the artefact. A person, not a team and not a process.
  • Recorded. Decisions and their reasons are written down as they happen, not reconstructed when someone asks.
  • Constrained. The hard lines are stated before the work starts: no number without a source, no claims about named individuals, stop when a conclusion needs a professional.
  • Evaluated. Anything reused, whether a prompt, a tool or a template, carries test cases with known answers and a failure threshold, and gets re-run when things change.
  • Disclosed. The reader is told what was checked and by whom, in a four-line block at the bottom: who authored it, who answers for it, what its limits are, and what backs it. There is one at the bottom of this article.

Four of the seven, the sourcing, the recording, the safety net and the re-checking, have no manual version at any real volume. You cannot diligence your way to them across everything you ship, which is why they belong in your systems rather than in your good intentions. The other three cost nothing but the habit.

One prompt, to run on the last thing you sent

Do not start by auditing your body of work. You will not, and if you tried, you would spend the effort evenly across things that do not matter. Take the last thing you sent. One document. Twenty minutes.

This is the prompt I use. It runs the checks that can run without fetching anything, and it tells you which claims need their sources opened. It is MIT licensed and free. No signup, nothing to fill in. Copy it off this page, or take it from the repo, and use it on whatever you like.

prompt
You are auditing a document I am about to send. Do not rewrite it, do not improve it, and do not comment on its quality.

First, find every load-bearing claim — numbers, findings, causal statements, regulatory assertions, and anything a reader would repeat. Restate each in one sentence and label it with exactly one of: **sourced** (the document names a source I could open, with a locator), **recalled** (asserted from memory, with no source or one too vague to open), or **inferred** (a conclusion the author drew). Note where a claim cites a write-up rather than the original, or a preprint rather than the published version.

Second, argue back. For the three claims the argument most depends on, make the strongest case that each one is wrong, and state what would have to be true for it to be false.

Third, group the sources into origins — ten write-ups of one press release are one source — and say how many independent origins actually support this document. Name any source that benefits from its own claim being believed.

Fourth, give a verdict: **defensible**, **thin**, or **not safe to send**, and the single most likely place this document is wrong.

Last, draft a four-line provenance block from what you found, not from how careful the author felt. Attribution: who authored it, and that AI tools assisted, in one sentence. Accountable: the named person who answers for it. Limitations: what is not backed, and how far each claim sits from its source — never empty. References: what backs it, where it lives, and how a reader can check it. Every line must either make a limit transparent or let the reader go and check something; cut any line that does neither.

Do not fetch anything. Where a claim cannot be trusted without opening its source, say so — that list, in order, is the author's homework.

The verdict is the part people skip and the part that earns its keep. Thin is the honest state of most documents, mine included, and it comes with a list of what to fix first. It is what turned a sentence I had quoted for weeks into a sentence I removed.

What one pass gives you

By the end of one pass you will have one document with its claims labelled, a verdict you can say out loud in a meeting, and a specific list of what to fix, in order. Start with the document you already feel slightly uneasy about. I'm fairly sure you thought of it while reading this.

The prompt is the half you can use this afternoon. The half that keeps working after you stop paying attention is in the same repo: the full framework write-up, the per-claim prompts, and three small tools:

  • sourced.py captures a source and hashes it, so months later you can ask whether the page changed underneath your sentence. It watches a single value on a page that moves, rather than the page, because on a live tracker the number is the finding and everything around it is noise.
  • sourced_status.py derives a claim's label from two things that actually vary: whether anything was retrieved, and how far the wording sits from the source. It exists because I fused those into one field, and the fused version could not describe a synthesis.
  • shared_claims.py finds claims used in more than one document and tells you when they have drifted apart. That one exists because a figure I corrected in two decks one morning was still wrong in two others that afternoon, and nothing connected them.

Both of the last two shipped with a bug in them first, and the repository says what the bugs were. One counted a placeholder as evidence and quietly promoted three unchecked claims. The other went briefly blind to the very error that prompted it. Trust them exactly as far as that record earns.

The rewards of checking are not abstract. The next time somebody asks you what you personally verified, you will have an answer, and the people who rely on your work will be standing on something that holds. If you run it on something and it finds you an eleven of your own, I would genuinely like to hear about it.

Attribution

Written by Andrew Ramsden. AI tools assisted research and drafting; all outputs verified.

Accountable

Andrew Ramsden.

Limitations

The Zittrain figures are from a 2014 study of legal citations and do not transfer directly to other kinds of document. Sources disagree on the size of the Deloitte refund, so this piece states the contract value and calls the refund partial rather than putting a number on it.

References

Each figure is quoted with a locator and was retrieved and hashed on 25 and 26 August 2026. The eleven corrections are held in three provenance sidecars beside the decks they correct, each carrying the verbatim quote, locator, retrieval date and content hash for every claim (unpublished raw data).

published