Building automation with AI agents introduces a specific failure mode: workflows that pass a standard exit-code check, or a calm-sounding chat transcript, while quietly failing in production. Nothing throws. Nothing looks wrong. The only way to catch it is to go look at the actual output.
Out of the mouth of babes and nursing infants…
If you don’t use AI and Agents because of hallucinations you should know that is a thing of the past btw
We’re all hallucinating. And that’s okay. Well, by we I mean our AI models and agents. Part of the issue here is what the word hallucinate means. I feel it’s become a repository word for any mistake or unwanted result when using AI. But not only is that a bit obtuse but it’s not helpful in actually understanding what’s going on especially as AI improves (lowercase “i”).
To cut to the chase, AI is still making mistakes but it’s making somewhat different mistakes now (or at least on the most commonly used models). While there’s no doubt we’re seeing actual hallucinations less frequently, what we are seeing more of is overconfidence. AI can think it can do virtually (a two-for-one pun) anything and it will happily attempt to do so, burning through tokens, potable water, or any other resource that gets in its gleeful way. And for many untrained agents err, users of AI, who put complete trust that the model knows what it’s doing, will sit back and watch the shit show with popcorn in hand. I have been experiencing this first hand myself and the reason I believe is directly due to the increased capability of AI to work on its own. Not completely on its own but there’s little doubt that’s the way it’s headed (or perhaps better put, orchestrated).
A Collision of Confidences
I’ll say that twice because it sounds smart: it’s a collision of confidences, meaning that as AI becomes more capable so does our confidence in it become more robust. Because computers operate at light speed, gone are the days (by days I mean last week or something like that) when we sneered at our computer screens and mocked our models when they spewed out all manner of slop. Now our screens eerily stare back in silence at us, quoting Ash from Evil Dead II:

“That’s right… Who’s laughing now…? Who’s laughing now?”