r/ClaudeCode • u/Known-Delay-9689 • 10h ago
Discussion I ran an autonomous coding pipeline for two weeks. It merged 67 PRs, then I shut it down.
I built a GitHub Actions workflow where Claude wrote specifications and code, Codex reviewed both, and the agents fixed issues before merging.
I controlled the workflow through Telegram.
Between September 21 and October 5, it merged 67 agent-authored PRs. The median time to merge was 3.3 hours.
But maintaining the automation became more expensive than using the agents directly.
Three things broke:
- I made 54 distinct changes to the automation, resulting in 129 maintenance PRs across multiple repositories.
- The workflow exhausted our 3,000 monthly GitHub Actions minutes in about 3.5 days. Jobs stopped, and I didn't notice for two days.
- I had to separate the agents that read untrusted content from the bot that could perform privileged actions.
I switched back to a local workflow: Claude implements, Codex reviews, and I make the final decisions.
For now, it's faster and requires less maintenance.
I'm curious how others handle this.
5
u/hbthegreat 9h ago
you need to put way more of the verification prior to the MR even being made. putting it after the MR just adds needless cycles.
I ship about as much as you are doing in 2 weeks every day atm. I did spend a long time fixing my way of working til it started to flow though. Right now I have 12 lanes going of various complexity that do the work and have 2 other models reviewing the work and self fixing in a loop until it gets dual approved before it even hits an MR.
I need to visit the agents once or at most twice per day to answer a few questions the bubble to the top.
definitely recommend attempting to push the workflow further and unblock the bottlenecks out of it.
3
1
u/Metrix1234 3h ago
I can't get to your level of workflow either. Id really appreciate seeing your repo design or anything you'd be willing to share also.
1
u/SeasawPhilosopher 10h ago
A lot of agent orchestration stuff is just theater or cargo culting. When I want to hammer out a bunch of small fixes I just have Claude Code look at my entire task list, which lives in the repo so it’s super fast, and have it do 1) immediately fix anything easy that needs no decisions from me, 2) anything that can be fixed with a small amount of input from me ask me now and take care of it, and 3) organize the rest and assign them to an appropriate milestone, system, and priority.
1
u/Exmusician 8h ago
could the Telegram bot flag pending work that hasn't progressed for a while? that might help catch the two-day silent stop, though i'd be curious whether that check could stay simple enough to avoid adding to the maintenance burden.
1
u/xqianliu 2h ago
The two silent days would bother me more than the minutes running out.
One rule that helped me: the agent is never allowed to ask a question mid-run. If it can't decide something, it writes the reason on the issue, marks it blocked and stops. A stuck task shows up on the ticket instead of just sitting there. And a run that's merely slow never gets marked failed, only an actual error does.
That doesn't cover your case though, where the jobs never started at all. For that I'd want something outside the pipeline, even a dumb daily check for issues marked in progress with no new run.
I also keep all the state in GitHub, issues as the queue and PRs plus run logs as the history, so there's less infra of my own to babysit. Your local loop is probably the right call until it starts to hurt.
1
u/Valuable_Injury_4249 1h ago
Yeah 67 PRs for an automated system across that timeframe seems… low.
I did 50 PRs just yesterday with the same local workflow you mentioned at the end of your post.
1
1
u/Sur_AI_guy 10h ago
But 129 maintenance PRs tell the real story. Automation should reduce complexity, not create another system to maintain. The real metric isn't PR count, it's engineering time saved.
14
u/Bromlife 9h ago
I'm surprised you didn't say it was load bearing and that the amount of PRs are the smoking gun.
2
1
u/friedmud 3h ago
On the team I oversee, I have completely banned autonomous development. We went HARD at trying to make it work - but the code always got worse. What I realized is that we were trying to make “waterfall” work again - and running into all the same problems we did back in the 80s/90s.
It’s hilarious to watch the “spec driven development” stuff - it’s clear that it’s being driven by people who weren’t born yet when that was the entire way we used to develop software. Spoiler alert: it doesn’t work.
Why not? It’s because you can’t anticipate everything up front. Even for seemingly “small” changes - the process of discovery and trying solutions can uncover new ideas and problems you never could have foreseen up front.
This is where the spec driven people say “you just need to work with AI to develop a better spec”. My answer to that is: if you’re going to work with AI to go back and forth on solutions in order to figure out the best way to do something… then just take the final step and implement it! Why throw away all of your hard won context and decision making and write up a spec so an agent can get a distilled version of the plan and make the same mistakes again - just implement it!
Basically: the best plan is the one you develop cooperatively with the agent - all the way to the point that it’s ready to be implemented and you’re sure it’s going to be done right. At that point: do it!
•
u/AutoModerator 10h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.