AI-Generated Evidence Is Only Caught By Accident


A Derbyshire Police case, and the reconstruction test most organisations have never run

A Derbyshire Police officer is under criminal investigation for allegedly using AI to create evidential material in a number of cases. The Crown Prosecution Service is now working with the force to review which cases may be affected. No arrests. The officer has been pulled off frontline duties while it is looked into.

It comes the same week PoliceAI launched, the new national centre meant to bring AI properly into policing. Backed by 75 million pounds of Home Office funding, it is meant to coordinate AI across all 43 forces in England and Wales. Its head told the press he had already had to step in and slow generative AI use down in places, because the accuracy bar was not being met.

Here is the part worth sitting with. He got caught. Somebody, somewhere in the process, looked at what was in front of them and did not accept it. That is the system working exactly as it is supposed to.

It is also the only reason we know about this one.


The Part That Worked

Somebody caught it. The honest assumption is that this was human review doing its job. A person in the loop looked at the material and felt something was off. Not a system flag. Not an automated check. A person, paying attention, willing to push back on something that read as plausible.

That is worth respecting. It is also worth being honest about what it actually proves.

Human review catching a fabrication tells you the reviewer was good that day. It does not tell you the process was good. Catching something like this depends on whether the person looking happens to have the expertise, the time, and enough suspicion in the moment. Swap the reviewer, change the workload, make the output slightly more convincing, and the same material probably sails through.

So the real question is not whether AI-generated “evidence” ever gets caught. It clearly does. This case proves that. The question is what happens in the cases where nobody had reason to doubt it. Where the output was confident, internally consistent, formatted the way evidence is supposed to look, and the person reviewing it had forty other things on their desk that day.

Nobody can answer that yet. Not in policing, not anywhere else AI is quietly producing the material that decisions get hung on. The only detection mechanism currently running is whether a human happens to notice. And nobody is measuring how often that fails.


Human Review Is a Hope, Not a Control

I have spent enough time on the governance side to know the difference between a control and a comfort. A control is something you can test, that produces evidence, that holds when the person operating it is average rather than excellent. A comfort is something that makes everyone feel covered until the day it is needed.

“There is a human in the loop” has become one of the great comforts of the AI era. It appears in policies, in vendor decks, in board papers. It sounds like a control. Most of the time it is a hope wearing a control’s clothes.

The Derbyshire case is human review at its best: it caught something. But notice what had to be true for it to work. The reviewer needed enough knowledge to feel the wrongness, enough time to act on it, and enough standing to challenge a colleague’s work. Take any one of those away and the same fabrication passes. A control you can disable by giving someone a busy week is not much of a control.


Not Just Policing

None of this is really about policing. The pattern is the same wherever AI is being used to produce the material that justifies a decision someone has already made.

An incident report where the timeline reads suspiciously clean. A root cause analysis that arrives at exactly the conclusion the project sponsor needed. A vendor due diligence file where the AI summary says the third party meets your security requirements, and nobody checks what it was actually summarising. A risk assessment that closes out a finding because the write-up sounds thorough, not because anyone tested whether it was true.

In every one of these, the document looks complete. It reads well. It has the right structure, the right tone, the language a real assessment would use. That is exactly the problem. Polish is not evidence. A document that looks like the output of careful work is not the same thing as a document that is the output of careful work, and AI has made the first one trivially easy to produce.

So here is the actual test, and it has nothing to do with how the document reads. Strip the AI-generated material out. What is left?

If what is left is original source data, timestamps, named people who made specific decisions and can be asked about them, you have evidence. If what is left is nothing, if the account only exists because the AI produced it and nobody can independently rebuild the same conclusion from the underlying material, you do not have evidence. You have a narrative that happens to be well-written.

That is the check. Not whether the AI output looks credible. Whether the conclusion survives without it.


Running the Test in Practice

It is not a paperwork exercise. Take one closed item from the last month. A report, an assessment, a decision write-up. Cover the AI-generated narrative and ask three questions of what remains.

First, is there a source underneath it that exists independently of the AI? A log, a dataset, a document, a measurement, a record someone else created. If the only source is the prompt, there is no source.

Second, is there a named person who made each material decision and could be asked to explain it without referring to the generated text? Accountability that only lives in the document is not accountability.

Third, if you handed the underlying material to a competent colleague who had never seen the AI version, would they reach the same conclusion? If the answer is no, or you are not sure, the conclusion was supplied by the tool, not supported by the facts.

Run that on a handful of files and you learn something uncomfortable but useful. You learn how much of your evidence base is real and how much is well-formatted assertion.


What Changes From Here

This case worked the way it worked because AI is still new enough to be imperfect, and so are the people using it to fabricate things. The fabrication had a flaw somewhere, or the reviewer had a reason to look twice. Right now, both of those are still likely.

They will not stay that way. The tools get better. The people misusing them learn what reads as convincing and what does not, what gets past a reviewer who has seen forty other files that week. The version of this that gets caught in a few years probably will not look like this one. It will look exactly like everything else in the file.

That is not a reason to panic about AI. It is a reason to ask whether your evidence stands on its own, or whether it is only standing because nobody has had a reason to look closely yet.

If you stripped the AI-generated material out of the last report, audit response, or due diligence file that crossed your desk, what would actually be left?


References

Discover more from Acceptable Risk (Documented)

Subscribe now to keep reading and get access to the full archive.

Continue reading