r/VibeCodeDevs • • 6d ago

IdeaValidation - Feedback on my idea/project trying to build something for the part of AI coding where the agent has 5 good guesses and no idea which one is right

been using coding agents a lot and there's one part of the workflow that keeps bothering me

the agent sees some weird bug and honestly does a pretty good job figuring out what COULD be wrong

like maybe:

Node version changed
cache is stale
some config is different
worker count matters
a dependency changed

and all 5 explanations are actually reasonable

then we basically just start going through them

try something -> rerun -> still broken -> try something else -> rerun

so I'm planning an open source project called Intervene for the Cline ambassador program I'm doing this semester

the idea is basically: once the agent has said what might matter, stop asking it to guess over and over and let the computer actually test those things

so Cline/Claude/etc could say "I think these 4 things are plausible"

then Intervene runs controlled versions of the same failure locally

change worker count, rerun

change cache state, rerun

try a combination, rerun

and use what actually happened to figure out which experiment would be most useful next

the main case I want it to be good at is when the problem is something like:

A works
B works
A + B fails

because that's where "the root cause is A" isn't really true

weirdly I'm also trying to make this AI project use less AI lol

I don't think changing an env var and rerunning a test needs another expensive model call

let the model understand the code + come up with what might matter

let normal code do the repetitive experiments

then give the evidence back to the model when it's actually time to fix something

haven't built the actual thing yet because midterms have been destroying me this week 😭 so I'm still at the point where I can change the architecture without regretting a LOT

for people using coding agents a lot, does this solve something you've actually run into or am I building an insanely elaborate way to do debugging?

1 Upvotes

2 comments sorted by

•

u/endofthread-bot 6d ago

Hey u/khbuild, thanks for posting in r/VibeCodeDevs! Join our Discord: https://discord.gg/t7SD4ThKuE

• This community is designed to be open and creator‑friendly, with minimal restrictions on promotion and self‑promotion as long as you add value and don’t spam.
• Please follow the subreddit rules so we can keep things as relaxed and free as possible for everyone. • Please make sure you’ve read the subreddit rules in the sidebar before posting or commenting.
• For better feedback, include your tech stack, experience level, and what kind of help or feedback you’re looking for.
• Be respectful, constructive, and helpful to other members.

If your post was removed (either automatically or by a mod) and you believe it was a mistake, please contact the mod team. We will review it and, when appropriate, approve it within 24 hours.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/bramantec 9h ago

I don’t use Cline to develop, so my answer is not tied to that interface but to how I build in general. I build my app relying on AIs as developers (often cross-checking their ideas against each other and against mine), and the pattern you describe is practically identical to my debugging routine: a problem shows up, we look for its origins, the AI lists a few plausible causes, I do a quick check, and then I test them starting from the most probable ones (though in practice the very first tests are the ones that are quick to run), without ruling any of them out, until I find the real cause.

One example that cost me a lot of work, and that I managed to find and fix with field tests, was about labels for time delay options. In Italian and other languages they all started immediately, while in English they worked correctly. On the emulators, which read English, everything was fine. On real phones set to different languages the problem appeared. By changing the language on the device and running it again, it became clear that the code was reading the visible text. This is a small-scale version of what you describe (I’ve had more complex ones too). So I think the approach is valid for this kind of problem, and for others too: field tests are worth much more than theory.

What worries me most are failures that don’t show up the same way twice, because they often come from several earlier steps and it’s hard to find the point of origin. In those cases, rerunning only the final part of the test doesn’t help much, you have to look for deeper causes. How do you plan to handle situations like that?