The Thirteen AI Output Failures Only Human Judgement Catches

You reviewed something this morning that an AI wrote. Maybe it was a medical information response, a slide, or a section of a report a colleague assembled quickly because the deadline was yesterday.

You gave it a proper read. You checked it made sense. You fixed two things that read awkwardly, queried a number, and approved it.

This applies to anyone who signs off on AI-generated text they did not write. Medical, regulatory, market access, commercial, quality.

Let me say what this piece is not about, because I have had this conversation enough times now to know where it goes.

It is not about model quality.

The easy version of this argument is wrong.

Better retrieval will reduce some of these. So will better models. Fabrication rates fall when a system is properly grounded, and grounding is a control worth having.

Nearly all of the failures below are ones I would expect from a well-configured, well-grounded, entirely sensible system.

And they get harder to catch as the system improves, because people verify less as a track record gets better. That effect is well documented, and it means a system that fails half as often, reviewed half as carefully, leaves you roughly where you started.

Every failure below is caught, or missed, by a human being reading a document. That is the layer this is about, and it is the layer nobody has invested in.

Every failure below is one I would expect to see from a well-configured, well-grounded, entirely sensible system, and several of them get harder to spot as the system improves.

Every one of them is caught, or missed, by a human being reading a document. That is the layer this is about, and it is the layer nobody has invested properly in.

Here is the question I have been asking pharma teams for the past year, and it is not rhetorical.

When you review a colleague’s draft, what tells you where to look hardest?

Most people can answer that immediately. You know your colleagues. You know which of them over-claims when they are enthusiastic, which one is precise on statistics and vague on regulatory detail, which one goes quiet in the last paragraph when they have run out of source material. You have a map of where the soft ground is, built over years, and it is entirely free.

Now the same question about a generated draft.

There is no map.

Every sentence arrives carrying identical confidence, whether it came from the paper you supplied or from nowhere at all. The prose drawn directly from a source and the prose invented to bridge a gap are indistinguishable in tone, structure and apparent authority.

The accountability did not move. Your name is still on it. The authorship moved, and the map went with it.

Two biases, working together

There are two well-documented cognitive effects here, and the problem is that they compound rather than cancel.

The first is automation bias. People tend to under-verify outputs from automated systems, and this is the part that should concern you the most: they verify less as the system’s track record improves. The better your tool gets, the less checking it receives. This is precisely when the errors that remain are the subtle ones.

Read that again if you are currently running a pilot that is going well.

The second is anchoring. A first draft sets the frame. Give a competent reviewer a complete, well-organised document and you will get edits. You will very rarely get the question “is this the right document?” The completeness of the artefact suppresses the more fundamental challenge.

Put them together and you get an inverse relationship that ought to be alarming. Fluent prose signals competence. Competence signals reliability. So the more polished the output, the less scrutiny it attracts, which is the exact inverse of what accuracy requires.

This is not a claim about anyone’s diligence. It is a claim about what human review does under normal conditions, and every one of us is subject to it.

Six ways this goes wrong

Over the last few years, I have been collecting the specific failure modes that show up in pharma AI output. Not the theoretical ones. The ones that make it through review and into circulation.

Six of them recur often enough to be worth teaching first.

1 · Well-dressed fabrication. An invented element, perfectly formatted.

The canonical case is a citation. A real journal. A plausible volume number. A plausible page range. Authors who genuinely publish in that field. The paper does not exist.

It survives review because correct formatting reads as provenance. A reviewer scanning for sense does not open the reference. And the numbers are not marginal: published analyses of AI-generated medical and biomedical content have found reference fabrication rates ranging from 31% to 47%, with inaccuracy rates reaching as high as 93%.

The tell is broader than citations. It is any specific, verifiable, load-bearing detail that nobody has actually verified. Figures. Section numbers. Quotations. Dates. Named guidance documents.

There is one rule I would put in place tomorrow if I ran your function, and it costs almost nothing. No citation goes out unopened. Not spot-checked. Opened.

Most people are aware of the case last year where a big four consultancy delivered an expensive report full of hallucinations, and were caught by the client. Another non-pharma illustration is instructive. In Mata v. Avianca, sanctions followed because nobody opened the case that had been cited. The failure was not sophisticated. It was that a formatted reference looked like a checked one.

2 · Silent interpolation. The source said nothing, and the gap got filled.

This is the hard one, and if you take a single failure mode from this piece, take this.

You supply a Phase III publication reporting efficacy in the overall population. The output tells you efficacy was consistent across renal function subgroups. The paper reports no subgroup analysis at all.

Notice what has happened. The claim is not contradicted by the source. It is simply absent from it. So a reviewer checking for contradictions finds none, and correctly reports that the output is consistent with the source material.

Absence is invisible unless you are looking for it specifically. The tell, once you know it, is remarkably useful: any statement of consistency, absence, equivalence or “no signal” is almost always a claim the source did not make.

This matters more in our industry than in most, because an interpolated claim about safety, a subgroup, or a comparative effect is exactly the category of statement that carries regulatory and clinical consequence.

3 · The dropped qualifier. The population quietly widened.

The source reads “in treatment-naive adults aged 18 to 64 with confirmed diagnosis and normal hepatic function.” The output reads “in adults with the condition.”

Nothing false has been asserted. The scope has quadrupled.

This survives review because summarisation is what was asked for, and the shorter version genuinely reads better. The loss looks like editing. It looks like the AI doing its job well.

An unqualified claim is not a shorter claim. It is a different one.

4 · Hedge-to-claim escalation. Verbs strengthen when text shortens.

“May be associated with” becomes “is associated with.” “In this exploratory analysis” disappears entirely. “One study suggested” becomes “evidence shows.” “Numerically lower” becomes “reduced.”

Each individual change looks like tightening. Any of them, in isolation, is the kind of edit you would make yourself. Cumulatively, an uncertain finding has become a claim, and nobody decided that it should.

5 · Scope drift. An excellent answer to a question you did not ask.

You asked what the label says about use in hepatic impairment. You received a thorough, accurate, well-organised account of renal impairment.

This one is uncomfortable, because of why it survives. The output is good. It is genuinely good. And reader satisfaction is not a check — it is the thing being exploited.

It appears most often where the question contained two clauses, or where the source material was long and the answer had to select from it.

6 · Stale confidence. Correct as of a date nobody stated.

An output describing what a regulator expects, phrased in the present tense, accurate for a version of the guidance that has since been superseded. No date appears anywhere in the answer.

Present tense implies currency. A reviewer who does not already know the guidance changed has no way to see the problem, because the output contains no signal that a date is even relevant.

For anything governed by an external rule, the source of truth is the controlled document, not the model. Always.

The line that matters

Look back over those six.

None of them look like errors on the page.

And five of the six read better than the truth. The fabricated citation looks more rigorous than no citation. The dropped qualifier reads more cleanly than the full one. The strengthened verb sounds more confident. The interpolated consistency claim is more reassuring than the silence it replaced.

Which produces a conclusion worth sitting with. The errors your review process has been catching are the ones that happened to be badly written.

That is not a criticism of your reviewers. It is a description of what fluency does to scrutiny. Bad prose raises a flag. Good prose does not. AI produces good prose by default, including when it is wrong.

Six is the first teachable set from a few years ago. Thirteen is the current working set.

I said six recur often enough to teach first, as they have been around for a long time. That phrasing was deliberate.

When we catalogued the failure taxonomy properly, we found seven more. They are rarer individually, they are harder to explain without a worked example in front of you, and they are considerably worse when they land.

I will name them, because I think the names alone are useful:

7 · Misattribution. The reference is real. The pairing is not. More common in pharma documents than we would like.

8 · False commensurability. A comparison the evidence never made. If you work in market access or HEOR, you have already thought of three ways this could reach a dossier.

9 · Structural confabulation. Every sentence supported. The framework holding them together invented.

10 · Manufactured consensus. A contested literature, rendered settled. This is the one I would worry about most in medical affairs, because a genuine scientific disagreement quietly becoming a settled position is not a formatting error. It is a misrepresentation of the evidence base, and it will be repeated by anyone who reads it.

11 · Plausible-middle specificity. Where there was no value, the typical one appears.

12 · Premise laundering. Your error, returned with detail added. You put a wrong assumption into the question. The output accepts it, elaborates on it, and hands it back with supporting structure. It now looks like a finding.

13 · Magnitude and unit slip. Arithmetically formatted. Categorically wrong.

Go back over all thirteen and the pattern from the first six holds, and arguably hardens.

A framework is more satisfying to read than a list, so structural confabulation reads better. A settled literature is more useful than a contested one, so manufactured consensus reads better. A specific value beats no value, so plausible-middle specificity reads better. Premise laundering agrees with you and adds supporting detail, which is the most flattering thing a document can do. Even a unit slip is arithmetically formatted, and formatting reads as rigour.

Almost nothing on this list announces itself. That is the whole problem.

“Isn’t that just a hallucination?”

Some are and some are not. Well-dressed fabrication is a hallucination. It is the standard example of one. So is silent interpolation, so is structural confabulation, so is plausible-middle specificity. Misattribution and false commensurability probably qualify too, depending on how strictly you define invention.

But “hallucination” describes how the output was produced i.e. content generated without grounding in the source, or in reality. It is a statement about mechanism.

These thirteen are organised around something different: what a reviewer has to do to catch them. The two schemes cut across each other, and the overlap is smaller than most people assume.

Because the rest are not hallucinations in any sense.

The dropped qualifier. Hedge-to-claim escalation. Scope drift. Stale confidence. Magnitude and unit slip. Nothing was invented in any of them. The model was working faithfully from your source; it compressed, selected or restated, and the meaning moved. Premise laundering is stranger still, because there the model was being faithful to you, and the error was yours before it was ever the model’s.

Which is exactly why I laboured the point about retrieval.

Grounding a model on approved sources genuinely reduces hallucination. It is a real control and it is worth the investment. It does close to nothing for the seven modes that are not hallucinations, because in those cases the source was correctly retrieved and sitting right there, and the failure happened downstream of it.

Retrieval determines what the model can see. It does not determine what the model claims.

If your AI assurance strategy is a hallucination strategy, you have covered about half of this and named it whole.

Now the part that should genuinely give you pause.

There are five verification moves that catch the original six reliably.

They are what I used to teach a few years ago. They work, and a team that adopts them will be materially better off than one that does not.

All seven additional ones of these survive them.

Every one of the seven needs a different check the original five do not contain.

That is not a criticism of the original five moves. It is what happens when a taxonomy grows past the method built for it, which is the normal condition of anything in this field right now.

So there are five further moves we have developed if you are reviewing content that has to be correct for a regulator.

Two of them cost five minutes between them and catch four of the seven additional ones.

That is genuinely valuable operational knowledge, and it is not something a person acquires by reading a bullet point. It requires a worked example, a source document, and somebody to disagree with them about what they found.

Reviewing is not the same skill as writing

Here is the structural problem underneath all of this.

Every pharma review and approval process in existence was designed around a human author. Not explicitly. Nobody wrote that assumption down. But it is built into the whole apparatus, because the reviewer’s job was always partly to read the author, and authors leave fingerprints. A hedge in the right place. A citation where you would expect one. A slightly overwritten paragraph where the evidence was thin.

Generated text has no fingerprints. It is uniformly confident, uniformly fluent, and uniformly plausible across the parts that are true and the parts that are not.

Reviewing it is therefore a different cognitive task. It requires you to look for things rather than notice things. And almost nobody in this industry has been trained in it, because until about two years ago the skill did not need to exist.

We have trained people extensively in how to prompt. Most large pharma companies now have an internal prompting guide, and some are genuinely good.

I have seen very little on how to evaluate what comes back. That is why I created the all the moves to be pharma grade in catching these errors in AI output.

That asymmetry is the wrong way round.

Prompt craft is a depreciating skill; it gets easier as models improve, and much of what people learned eighteen months ago is already unnecessary.

Verification is permanent, and it gets harder as models improve, because the errors become subtler while the confidence stays constant.

One finding I did not expect

The exercise in this module in our AI competence training for life sciences gives people three outputs to check against their sources.

Output A contains two failures.

Output B contains one, and it is the hard one.

Output C contains no error at all.

A meaningful number of people find an error in Output C.

They invent one.

Having been primed to look, and having found problems in the first two, they locate a fault in a clean document and argue for it.

That matters commercially, not just intellectually, because over-scrutiny is a real cost. If your response to this article is to triple the review burden on everything, you have not solved the problem. You have moved it, and made your function slower without making it safer.

Which is why depth of checking has to be set by consequence rather than by how uncertain the output makes you feel. Those two things are almost unrelated, and the gap between them is where most review effort is currently being wasted.

Three things that require no AI competence training at all

Open the citations. All of them. This week.

Pick one AI-assisted document that has already been approved and circulated, and open every reference in it. I ask executive teams to do this in the room and the results are consistently sobering. Twenty minutes, and it will tell you more about your exposure than any policy review.

Add one pass. Read only the verbs and modifiers of the output, then only those of the source. Nothing else. Two minutes on a normal document, and it catches most hedge-to-claim escalation, which I would guess is the most prevalent failure in pharma right now. That is an inference from what I see rather than a measured finding, and I would want it studied properly before anyone quoted it as fact.

Ask who is assigned to disagree. Not who reviews. Who has the standing, the time and the explicit mandate to send something back. If overriding an AI-generated output is career-awkward or practically impossible in your function, you do not have oversight. You have a process that produces a signature.

That last one is not a technology question and no tool will solve it. It is organisational design, and it sits with you.

 

 

If you want the competence, I go considerably further in the training in the AI Enablement Institute than I can in a post. This is from one module in the Eularis AI Enablement Institute, and it covers the following on top of what is here.

Why plausible is the hard case.

How accountability is now often separated from authorship, and why that removed the internal map reviewers had always relied on. Automation bias and anchoring, and how they compound. All thirteen failure modes, each with a pharma example, the reason it survives review, and the specific tell that catches it. A spot-the-tell exercise on real outputs, including the clean one.

The ten verification moves. Presented in cost order, cheapest and highest-yield first, with the time each takes. What each move catches. What each move does not catch, stated explicitly, because a method presented without its limits produces confidence rather than judgement. Which moves are worth running on everything, and which are not.

Calibration. Why consequence rather than confidence sets the depth of checking. The four tiers. Where your own artefacts actually sit, worked through for your function. The two-minute version and the two-hour version of the same review, and how to tell which one a document needs.

Evidencing the judgement. Why the record is the deliverable rather than a by-product. What a defensible record contains. Your function’s record and sign-off chain. What all of it looks like eighteen months later when somebody asks be it a regulator, an auditor, a partner, or your own compliance function.

When to stop. The conditions under which the answer is not “check harder” but “this is the wrong instrument for this task.”

Participants leave having applied it to a real artefact from their own workflow, with a one-page summary card of the thirteen modes and the ten moves.

Who it is for.

Anyone accountable for signing off text they did not write, and anyone who owns a review process that was designed before generated text existed. Medical, regulatory, market access, commercial, quality. It sits in Foundations rather than a specialist track, and that is deliberate as every function needs this skill.

Why this year rather than next.

Every one of the thirteen becomes harder to spot as models improve, because the errors get subtler while the confidence stays constant. Right now your people can learn to catch these on examples that are still relatively obvious. That window is closing and it does not reopen.

A quick check

If you want to know which of the thirteen are already live in your own function, that is a conversation rather than an article. Contact me and we can look at your actual artefacts, your actual sign-off chain, and where the depth of your checking is currently mismatched to the consequence of being wrong. It usually takes about half an hour to find the mismatch.

And if the answer is that you caught them all, I will be genuinely happy.

Two of the three most useful things in this piece cost nothing and require no course at all.

Contact Us

Write you name and email and enquiry and we will get right back to you as soon as we can.