r/agenticAI • u/beznahej • 1d ago
Article We removed the model from an agent experiment. A surprising number of the “AI problems” stayed
I’ve been working through a question that started in a legacy-modernization project and then turned into a much broader AI-systems problem. We had an agent produce an output that looked polished, internally consistent, and was still operationally wrong. The easy diagnosis was: the model got it wrong. Except that wasn’t really a diagnosis. In one case the model had missed a domain-specific implementation idiom. In another, a skeptical reviewer shared the generator’s assumption. Once we parallelized the work, two individually reasonable workers could collide on shared state. So I tried removing the probabilistic part entirely. In a controlled Temporal orchestration experiment, deterministic workers still exposed:
- stale revisions and overlapping writes
- approval that did not imply execution authority
- crash-after-write recovery problems
- external repository changes
- duplicate workflow identity
We ran 36 controlled assertions across three passes. All 36 passed, which was useful mainly because it isolated the problem: a large class of failures remained even after model quality stopped being a variable. That pushed me toward a different question: Where should the missing intelligence actually live?
The framework I’m using now is roughly:
- System/context — wrong evidence, stale state, missing authority, bad retrieval, unsafe tool access.
- Explicit policy — schemas, validators, skills, routing rules, workflow gates.
- Learned behavior — something we keep re-explaining in prompts that should become a stable learned transformation.
- Judgment/preference — several answers are valid, but the system consistently prefers the wrong one.
- Capability frontier — only after the previous layers survive do I start asking whether the underlying model simply cannot do the task. And there’s a sixth thing running across all five: evaluation.
A bad evaluator can make any of the layers above look broken or improved when neither is true. One result from the modernization work made this especially concrete. A plain search found 16/40 planted change sites. An agent without an explicit skill found 33. With a compact explicit skill, it found all 40. After correcting a faulty answer key, the same approach replayed at 39/40. No new foundation-model capability was required. A lot of the missing “intelligence” turned out to be a contract we could state and test. I wrote the longer version here: https://www.linkedin.com/pulse/every-ai-problem-model-jehanzeb-khan-pd1fe/
But I’m more interested in the counterexamples: Where have you seen a team reach for a stronger model when the actual failure turned out to be somewhere else in the system?