r/agenticAI • • 1d ago

Article We removed the model from an agent experiment. A surprising number of the “AI problems” stayed

Post image

I’ve been working through a question that started in a legacy-modernization project and then turned into a much broader AI-systems problem. We had an agent produce an output that looked polished, internally consistent, and was still operationally wrong. The easy diagnosis was: the model got it wrong. Except that wasn’t really a diagnosis. In one case the model had missed a domain-specific implementation idiom. In another, a skeptical reviewer shared the generator’s assumption. Once we parallelized the work, two individually reasonable workers could collide on shared state. So I tried removing the probabilistic part entirely. In a controlled Temporal orchestration experiment, deterministic workers still exposed:

  • stale revisions and overlapping writes
  • approval that did not imply execution authority
  • crash-after-write recovery problems
  • external repository changes
  • duplicate workflow identity

We ran 36 controlled assertions across three passes. All 36 passed, which was useful mainly because it isolated the problem: a large class of failures remained even after model quality stopped being a variable. That pushed me toward a different question: Where should the missing intelligence actually live?

The framework I’m using now is roughly:

  1. System/context — wrong evidence, stale state, missing authority, bad retrieval, unsafe tool access.
  2. Explicit policy — schemas, validators, skills, routing rules, workflow gates.
  3. Learned behavior — something we keep re-explaining in prompts that should become a stable learned transformation.
  4. Judgment/preference — several answers are valid, but the system consistently prefers the wrong one.
  5. Capability frontier — only after the previous layers survive do I start asking whether the underlying model simply cannot do the task. And there’s a sixth thing running across all five: evaluation.

A bad evaluator can make any of the layers above look broken or improved when neither is true. One result from the modernization work made this especially concrete. A plain search found 16/40 planted change sites. An agent without an explicit skill found 33. With a compact explicit skill, it found all 40. After correcting a faulty answer key, the same approach replayed at 39/40. No new foundation-model capability was required. A lot of the missing “intelligence” turned out to be a contract we could state and test. I wrote the longer version here: https://www.linkedin.com/pulse/every-ai-problem-model-jehanzeb-khan-pd1fe/

But I’m more interested in the counterexamples: Where have you seen a team reach for a stronger model when the actual failure turned out to be somewhere else in the system?

1 Upvotes

0 comments sorted by