In August 2026 we generated six internal documents — a decision memo, a client permission letter, an operational runbook, a directory pack, a task list and a set of ticket bodies. All six were produced carefully, with access to our own records, by a process designed to be thorough.
Then we ran a second pass whose only instruction was: try to prove each of these wrong.
None of the six survived. Between them they produced 29 problems serious enough to stop publication. This is what those problems looked like, because the pattern is more useful than the count.
The failures were not errors of fact so much as errors of provenance
Almost nothing was invented outright. That is the part worth internalising: the documents were not hallucinating, they were misattributing.
The most instructive case involved a figure in our own records marked "re-verified". That note meant one specific thing — the number reproduces from the source it claims. Two separate documents read it as a general clearance and used the figure as a headline proof point.
It was not safe to use. Our own earlier audit had recorded that the underlying metric counted the wrong events. The number was real; what the sentence around it implied was not. And the note that caused the confusion was one we had written ourselves, months earlier, for our own convenience.
The averages problem
A second document led with an average drawn from several accounts we manage. The figure was arithmetically correct. It also silently excluded accounts that would have moved it substantially in the other direction, and no inclusion rule was stated anywhere, so nobody reading it could have reproduced it.
There was no deception involved. The exclusions were reasonable individually. But an average whose sample frame is undisclosed is not a measurement, it is an assertion wearing a measurement's clothes — and it is the single most common way a document is wrong while every individual sentence in it is true.
Confidentiality did not travel with the content
This one changed a process rather than a document. One memo was correctly marked internal-only at the top. It also contained a block formatted for pasting into a ticket, and that block — the part actually designed to travel — carried no marking at all, and reproduced several named clients' account data in a single paragraph.
The lesson generalises well beyond AI-assisted work: a confidentiality marking on a document does not protect the payload inside it. If part of a document is designed to be copied elsewhere, the warning has to be on that part.
Four of the six documents leaked client-identifying performance data into a surface built to be shared. Every one of them was written by a process that had read the rule against doing exactly that.
Why proofreading does not find these
Read any of the above in context and it reads as competent work. The prose is clear, the structure is sound, the citations point at documents that exist. A proofreader checks whether the writing is good and whether the claims are plausible, and by both measures everything passed.
What surfaced the problems was a different instruction. Not "review this", which invites agreement, but "here is a specific claim — find the reason it is false, and default to rejecting it if you cannot resolve the question".
That framing matters more than the tooling. A reviewer asked to check will confirm. A reviewer asked to refute will go and open the source.
What we do differently now
- A claim is verified against the thing it cites, not against a note about the thing it cites. If the note is ambiguous, the note is the bug — fix it once rather than re-reading it correctly forever.
- Any aggregate states its sample frame inline: which accounts, which window, what was excluded and why. An average without a frame does not ship.
- Confidentiality markings go on the payload, not the document. If a section is designed to be pasted somewhere, it carries its own warning.
- "Re-verified" is never a status on its own. A verification note records which property of a claim was re-checked.
- The reviewer is not the writer, and is briefed to refute rather than to confirm.
The honest summary
The generation step was fast, cheap and good. The output was well-organised, well-written, and would have embarrassed us in five places and exposed a client in a sixth.
That is not an argument against using these tools — we use them constantly, including for the boring audit work that finds problems like these. It is an argument that the review step is now the expensive part of the job, and that pretending otherwise is how an agency ships something it cannot defend.
If your agency is producing more deliverables than it used to and the review process has not changed, that is the thing worth asking about.
Common questions
Does this mean AI-generated content is unreliable?
It means it is unreliable in a specific and predictable way: confident presentation of claims whose provenance is weaker than it appears. That pattern is checkable, which makes it manageable. The mistake is assuming that fluent, well-structured output has already been checked.
Can't you just ask the model to check its own work?
Partially, and it helps more when the checking instruction is adversarial and the checker does not have the original reasoning in front of it. A reviewer told the conclusion is probably right tends to find reasons it is right. Independence matters as much for models as for people.
How much review is enough?
Scale it to the cost of being wrong. An internal note needs a glance. Anything that goes to a client, gets published, or informs a spending decision deserves a separate pass that tries to break each specific claim. The claims that most deserve attention are the ones doing the most persuasive work.
What is the single highest-value check?
Take the most impressive number in the document and trace it to its source. Not the source it cites — the actual data. In our experience that one check finds more than a full read-through, because the most persuasive figure is the one most likely to have travelled furthest from its origin.
Isn't this just normal editorial quality control?
In principle yes, and that is rather the point. What changed is the volume of material arriving for review and the fact that its surface quality no longer correlates with whether it was checked. The discipline is old; the reason it is now load-bearing is new.
Want us to check your site?
We will look for the same faults on your website and tell you what we find, whether or not you work with us.
Book a free consultation