I built an AI system this year whose only job is to catch compliance mistakes before they ship.
It works well. Almost too well, until it doesn’t.
Nothing, not worse
Give it a case close to something it has seen before, and it catches things a tired human would miss. Give it something genuinely new, no comparable case anywhere in its training, and it doesn’t get worse. It gets nothing. There is simply no ground to stand on yet.
I watched this happen enough times to stop calling it a bug. It defers to me exactly when the case has no precedent behind it. Not most of the time. Every time.
So I built that deferral into the system on purpose, instead of treating it as a gap to be trained away.
What a deferral is worth
Here is the part that doesn’t make it into most AI conversations. A deferral, on its own, is nothing. Escalating a hard call to a human and then losing the answer the moment the conversation ends is just a slower way of staying stuck.
What makes the deferral worth having is what happens after: the ruling gets written down, permanently, with its own identifier, tied to the exact question it answered and the exact reasoning that came with it. The next time something close to that case appears, the system doesn’t ask me again. It points back to the ruling by name.
That record is not a changelog nobody reads. It’s a growing index the system checks before it hands me anything. It now runs past fifty entries, and every one of them is a decision I actually made, on a case that actually happened, kept in a form the system can search rather than a form only I remember.
The floor rises, the ceiling doesn’t move
That’s the part I underestimated when I built it. I assumed the deferrals would taper off as the record grew. They haven’t. What’s changed is what they’re about.
Early on, I was getting pulled in on calls that, in hindsight, weren’t that hard. Common enough shapes that a wider precedent record would have covered them eventually. Those calls have mostly stopped reaching me. The record absorbed them.
What still reaches me are the cases with no shape yet. And there is no version of this where that category empties out, because “genuinely new” isn’t a backlog you clear. It’s a moving target by definition. New situations keep showing up faster than any system can pre-train against all of them, because that is what genuinely new means.
So the loop never closes. It compresses. The floor of what needs a human keeps rising, but the ceiling never disappears, because the ceiling is just wherever precedent runs out, and precedent always runs out somewhere.
That’s not a temporary gap while the technology catches up. I don’t think it closes at the next model, or the one after that. It’s a permanent seat at the table for a human who has actually made the call before, sitting exactly where the record stops and the genuinely new begins.
I didn’t build this system to prove that point. I built it to stop making the same compliance mistake twice. The seat that never empties was just what was left standing once I looked at what the system was actually doing, case by case, over enough weeks to see the pattern instead of the headline.