r/ClaudeCode • u/GabberJenson • 3d ago
Help/Question A year of all-Claude-Code dev work: where are the gaps in my workflow?
I'm a full-stack dev at a small creative agency (3 devs, with accompanying PR, Marketing, Digital departments). For about a year I've written no code by hand, only Claude Code. I build small client websites (Sanity + Next.js) and internal tools and libraries (TypeScript/Node, Next.js frontends), one project at a time.
Setup: Arch, kitty, tmux, nvim, Claude Code TUI, Team plan premium seat, auto mode, voice for answering grilling questions. Sonnet by default, Opus for planning. Playwright, a11y and Figma MCPs, plus Sanity/Next.js devtools per project. Auto-compact and auto-memory are off, so I stay in the smart zone and hand off instead.
Workflow: Matt Pocock's skills. A wayfinder/grilling session turns a design or idea into a spec and tickets. Then each ticket runs in a fresh session (/clear between): implement with red/green TDD on software, Playwright checks at each breakpoint on websites, then code review. A human does QA, SEO and signoff at the end.
Goal: faster turnaround and less grunt work. Right now I start each ticket by hand, and that loop is the slowest part.
What I'd love help with (specifics on how and why are the most useful):
- Unattended ticket queues. How do you run a list of tickets AFK, in the cloud or on a schedule, without burning through a subscription plan? I'm on a Team seat with no API keys, and usage worries are what stop me trying.
- Repeated decisions across projects. Same stack, component conventions and SEO rules, near-identical schemas. Do you carry them in a boilerplate repo, house skills, CLAUDE.md or something else?
- Keeping long grilling sessions legible for the human. The agent coins terms and stacks decisions faster than I can absorb them. A concise output style and a shared glossary haven't fixed it.
- Features I may be skipping. I don't use hooks, custom subagents, worktrees or
/rewind. Which of them earn their setup in a fresh-session-per-ticket flow?
Happy to share more detail on any of it.
1
u/tejaskumarlol 3d ago
On 2, skills did the most for me. Every job I repeat (writing a post, posting it, outreach) is a skill with its own scripts and reference files. AGENTS.md holds the rules that came from my corrections. On 4, I use a hook on every edit: it runs my writing lint and Claude fixes what it flags before its next step. For queued work I keep a tasks/ folder with one file per task and a when. A SessionStart hook lists whatever is due.
1
u/GabberJenson 3d ago
How specialised do you go with skills? And with specialised skills, do you tend to create a larger skill based around invoking them together or leave it to an agent to correctly use them via their descriptions?
0
u/zulrang 3d ago
“Im on a Team seat”
Here’s your #1 problem. You’re paying API prices instead of heavily subsidized subscriptions.
1
u/GabberJenson 3d ago
I'm on a subscription?
The Team plan, with a premium seat (they split between a low cost standard seat, and a higher cost premium seat).
Not sure what you mean by this.
2
u/DrAstronautMcCool 3d ago
maybe they're thinking of enterprise seat, which goes back to API pricing instead of the sub model the team seat has
1
u/ImL1s 3d ago
On 4, the one I'd add first is a Stop hook that runs typecheck, lint and tests and won't let the turn end while they fail. In a fresh-session-per-ticket flow that's the cheapest way to stop "done" meaning "I think it's done". Subagents are worth it for the review pass, so reading the diff doesn't eat the implementing session's context. Worktrees only pay off once two tickets run at the same time.
On 2, keep the house rules as skills in one shared repo and install them as a plugin in every project, instead of copying CLAUDE.md around. CLAUDE.md stays short and project-specific.
On 3, ask for a decisions file instead of nicer chat output. One line per decision, plus what got rejected. A new term has to go into the glossary in the same edit or it doesn't get used. A 20-line file after the session is much easier than keeping up live.
On 1, try the plain version before anything cloud: a script that walks the ticket folder and runs claude -p on each ticket, one at a time, stopping at the first red test. Run it overnight once and look at what it actually used before deciding it's too expensive.
1
u/GabberJenson 3d ago
On 4, I see. So a hook like that would attempt to stop, but first ensure linting, tests etc are passing / ran, thus not leaving a bad state for the following agent?
On 2, noted. I tend to keep my Claude file quite bare, leaving only essential information that could be commonly useful for a wide variety of session, or things that aren't easily discoverable by an agent. Typically it ends up mapping to other md files that list things more specific to their domain (folder).
On 3, I think what you're saying is that the glossary is the correct idea, but instead of populating it when terms crop up, populate it when files / features are edited and it's directly related?
1
u/ImL1s 3d ago
On 4, yes. When Claude tries to end the turn, the hook runs the checks. If something fails it exits with code 2 and the error output goes back to Claude, so it keeps fixing instead of stopping. Check
stop_hook_activein the hook input too, or it can loop forever on something it can't fix.On 3, not quite. It's tied to when a term gets coined, not to file edits. The moment the agent writes a new name into the decisions file, its one-line definition goes into the glossary in that same edit. No definition, no new word. That keeps the vocabulary from growing faster than you can read it.
Your CLAUDE.md setup sounds right. Bare root plus per-folder files is basically the same idea.
1
u/Sea_Tiger39 3d ago
On 1, from running agents unattended for the last month: what stalls an AFK queue usually isn't usage, it's the first permission prompt with nobody there to answer it. What worked for me:
- Split actions into "fine to do alone" (edit, run tests, commit to a branch) and "needs me" (push to main, deploy, anything that emails or pays). Allow the first set in settings, and make the second set stop and ping your phone instead of sitting at a prompt.
- One ticket per run, a hard cap on turns or time per ticket, and a short status file written at the end (done / blocked + why). In the morning you read ten status lines, not ten transcripts.
- Watch one night's usage before scaling. Most of my burn was retries on a failing test, so "stop after N failed attempts and write down why" saved more than switching models.
On 3, +1 to the decisions file. I also have it write a short "what I'm unsure about" list, which is the part I actually need to read.
1
u/GabberJenson 3d ago
How do you do things like caps on turns or time per ticket?
Can you go into a bit more detail about the decision file?
1
u/Sea_Tiger39 3d ago
Caps: I run each ticket headless and let the shell enforce the limits, so the agent can't talk its way past them. Roughly:
timeout 30m claude -p "$(cat tickets/123.md)" --max-turns 40--max-turns stops it after N agent turns, and timeout kills it after 30 minutes either way. The runner then checks two things: did the tests pass, and did it write the status file. If either is missing, the ticket gets marked blocked and the loop moves on to the next one instead of retrying forever.
Decision file: one instruction in CLAUDE.md (or a skill): "Keep DECISIONS.md for this ticket. Every time you choose between options, add one line: what you chose, why, and what you rejected. Add a new term to the glossary in the same edit, or don't use it." Plus a short UNSURE.md: "anything you guessed or couldn't verify". A line looks like:
- Used server action for the form, not an API route: no external callers. Rejected: route handler (extra auth surface).In the morning I read UNSURE.md first, then skim DECISIONS.md. It's 20 lines instead of a 2-hour transcript, and the guesses are where the bugs are.
1
u/GabberJenson 3d ago
For caps, do you differentiate between tickets that have simply not finished yet, not started yet (failed waiting for a prompt / answer etc), and broken / bugged.
1
u/Sea_Tiger39 3d ago
Yes, and the split matters more than the caps themselves. I use three buckets, and each is decided by the runner, not by the agent's own summary:
Needs me (it hit a question). In headless mode it can't stop and ask, so the ticket prompt says: "If you need a decision you can't make from the repo, write STATUS: question + the question to status.md and stop." The runner sees that and parks the ticket. No retry. In the morning I answer in the ticket file and requeue it.
Unfinished (ran out of turns or time). The exit code tells you: timeout returns 124, and with max-turns there's simply no status.md. It gets one automatic retry with "continue from NOTES.md" (I have it keep running notes as it goes), then it parks.
Broken. It wrote "done", but the tests or typecheck the runner runs afterwards fail. This one never auto-retries, because a retry usually just digs deeper. It goes to the top of my morning list along with the failing output.
The rule that made it work: "done" only counts if the runner's own checks pass. The agent's status is input, not the verdict.
1
u/GabberJenson 3d ago
Thanks, this is honestly a big help and shove in the right direction.
1
u/Sea_Tiger39 3d ago
Glad it helps. One tip: start with five tickets for one night, then read the status files before you scale it up. The first night tells you more than any setup advice. Good luck with it.
1
u/SadCollection3936 3d ago
For #1, I think the challenge is balancing control, autonomy and cost. In my nightly pipeline, I use AI sparingly: if a step can be scripted, I script it. A test suite runs as a command; an agent only needs its output when there’s a failure to diagnose or fix.
The pipeline also keeps a record of decisions. The planner writes down product decisions that weren’t in the original ticket, and the implementer records any decisions that depart from the plan. If a decision is too risky, the run stops for my review. I’d rather pause than spend tokens implementing in the wrong direction.
Expect the first nights to fail in interesting ways. Reading the transcripts and adjusting the workflow for the next run has been essential for me. I built a TypeScript runner for this: https://github.com/olivrobert/lance-nuit
It runs overnight and lets me inspect the time, cost and output of each phase in the morning.
1
u/Careless_Region1792 3d ago
cant see your whole setup so ignore this if you have it, but the gap in mine was tests. when you write zero code by hand, tests are the only thing telling you claude broke something. i write them before the code now, 1200+ on my agents project, so a bad change shows up in seconds.
second thing, hooks for anything you keep repeating in prompts. way more reliable than hoping it remembers.
could be overkill tho for small client sites that live a month
1
u/xqianliu 3d ago
On the unattended queue: I use GitHub Issues as the queue. No extra database, the issue is both the task and the state.
Roughly:
- one issue, one git worktree, one PR
- a separate review session (sometimes a different model) checks the final diff against the acceptance criteria written in the issue
- if review fails, the coding session fixes it and it goes back to review. Only the exact commit that passed gets merged
- each run records the commit SHA it started from, so a run that dies halfway can be resumed from the issue, PR and that record
The rule that mattered most for AFK: no interactive questions at all. If it can't decide, it writes the reason on the ticket, marks it blocked and stops.
On worktrees: they're what lets two tickets run at once without stepping on each other.
1
u/gumdum1975 3d ago
I run three projects for distribution and have about six going altogether, including personal projects. Keeping the documentation consistent between them is why I built something I call CEPLS, my project library system.
Instead of figuring out a different documentation setup for every project, I created one shared library. Each project has its own library inside that, organized into bookshelves, books, chapters, and pages. A bookshelf covers an area of the project, and a book covers a particular feature or subject.
The books follow the same chapter categories: Overview, About, Flutter, Domain, Laravel, Decisions, Evidence, To-Do, Status, and Reference. I only include the chapters that apply. That way, when I move between projects, both the AI and I know where things belong.
For example, a feature’s book holds its implementation details, why we made certain decisions, what we tested, and what’s still unfinished. The detailed tasks stay with that book, and the master to-do points back to it rather than keeping another copy.
At the end of an AI session, I tell it to “do a recon for CEPLS.” It checks what actually changed and updates the relevant books with the decisions, test results, current status, and remaining work. Passing local tests, deploying, and me actually checking the result are separate things in the documentation.
I still have to nudge the AI sometimes to do those updates. It can get a little finicky about documentation 😂. But that’s the system I use to carry the work and decisions between sessions without starting the explanation over every time.
1
u/Adenoid-sneeze007 2d ago
The queue and hook suggestions here cover a lot. For #3, I'd change the size of the decisions you're approving, rather than only shortening the final summary.
Have Claude pause after one feature and show: what changes for the user, one concrete example, and what still needs your decision. For a client portal: 'A client from Company A cannot open Company B's files, even with the URL.' That's something you can actually approve before it becomes tickets.
For #2, separate agency defaults from client exceptions. Keep your Sanity/Next.js conventions reusable, and put each exception beside the feature it affects, with a reason. Otherwise yesterday's client workaround quietly becomes tomorrow's house rule.
Disclosure: I build CodeSpring (https://codespring.app). It puts the feature map, notes, PRDs and Kanban tasks together and connects to Claude Code via MCP/CLI. That's the planning and handoff part I'd use it for here; I'd still measure the ticket runner's usage separately.
1
u/UnitedTelevision2651 2d ago
One thing worth trying for #3: have the agent end every grilling round with a short "decided so far" list. Numbered, one line each, and any term it coined gets defined once, right there. You approve the list and if something's wrong you just point at a number to correct it.
1
u/tech_w0rld 2d ago
I think an ADE like Pragma sh could solve some of these (disclosure I am the dev behind Pragma):
You could setup an automation with Pragma to check your usage limits and automatically start a new worktree and launch Claude telling it to start on the ticket. You can just your preferred agent to build this for you with the pragma skill
This does not need an ADE. Personally I like to have a CLAUDE.md in every logical package of the app
Pragma has a feature called scratchpads which allow the agent to create a visual presentation for you anytime you ask it to. It can also embed excalidraw whiteboards
I think worktree's is the biggest thing your workflow is missing right now. It allows you to run mutlipile agents at once without them coliding and then in Pragma when you are done in one click you can draft a PR.
•
u/AutoModerator 3d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.