I have been on both sides of a system like this. On one side are the people building it. On the other is the person asked whether the output is any good. For a while I was that second person.
The system automated a review that would normally be carried out by a subject matter expert. This was not an advisory output that somebody would double check later. It fed a formal process, and people downstream relied on what it produced.
It was built to specification, and it was generating solid, well-structured reviews that passed every test case defined for it. To the people who built it, the system was working as designed. To an expert in the field, it was not. For AI, “working as designed” and “correct” are two different things, and the build team cannot test its way to the difference.
What I saw
I reviewed the output on a regular basis, to make sure each software update was improving the system rather than breaking it. On several occasions the output was inaccurate for someone who knew the subject in detail. It was corrected. These gaps were invisible to the people building it, and it is worth being precise about why, because it was not a matter of diligence.
Conventional software tends to fail in ways that announce themselves. A total does not reconcile, a request returns an error, a field comes back empty. There is a signal, and anyone paying attention can act on it without knowing much about the subject matter. A system that produces written analysis does not behave that way. Its output is well structured, fluent and confident even when it is wrong, because producing something that reads correctly is close to what it was optimized to do. An inaccurate review does not look like a defect. It looks exactly like a competent review.
That leaves a non-expert reviewer with nothing to notice. The engineers could confirm that the system ran, that it produced output in the expected shape, and that it handled the cases there were tests for. What they could not confirm was whether the analysis was correct, and no amount of additional care would have changed that, because answering the question requires knowing the subject well enough to produce the analysis yourself. If someone on the build team could have done that, the system would not have been needed.
The people best placed to build a system of this kind are, by definition, the people least equipped to judge what it produces.
But surely monitoring solves this
The development environment was a solid one. It included the tooling needed to observe and report on the performance of the system and the underlying models, and that tooling did its job. The observation dashboard produced the intended results and gave us insight that was pivotal in improving the output. Drift detection, evaluation sets and output monitoring do work, and in most situations, they are the right answer. They did not catch this, and the reason is not that they had been configured badly.
What was observed was the wrong thing to observe. You observe the right thing when you know exactly what needs to be observed, and in this case that was the content of the review the system produced. To judge that, you must be an expert in the subject.
What actually closes the gap
A dedicated subject matter expert is not a nice addition to a project of this kind. It is the control that makes the output trustworthy, and it must be treated as an ongoing responsibility rather than a review at launch. That person observes and reports continuously, and where an update moves the output away from what it should be, they roll it back or reject it.
Three things make that real. Someone must be accountable by name. They must check at a set frequency, not when somebody remembers to ask. And they must hold the authority to stop the system when it starts behaving differently from its intended purpose. Without the third, the first two are decoration.
Luck is not a control
In that case the problem was caught, and it was caught by “luck of team design”. The person who could see it happened to be the person who could act on it. That is not a control embedded in how the organization was built. A clear governance system is what turns that luck into a control: named accountability, a defined frequency of measurement, and the authority to correct a deviation once it is found.
There is a body of standards work that describes exactly this machinery, and I will come to it in the pieces that follow. None of it will help an organization that still believes a green test suite is the same thing as a correct answer.