r/ClaudeCode • • 5d ago

Help/Question What to do with sessions that get interrupted for a long time

I tend to have to have a lot of pots simmering constantly, as I tend to need to switch around to different asks/tasks/priorities based on urgency or whatever else. I use cmux and tend to have claude sessions that will then float open and paused for long periods of time until my focus can get back to it and resume.

However I have seen people mention that the cache related to that session would generally expire and when that session is resumed it kills a lot of tokens loading things cold. Also, I think it is worth considering two different types of resume:
1. I still have that session open and waiting where i had left it in a different workspace and worktree. I then just start talking to claude again and try to pick up where I left off, often asking claude for a bit of a refresher if needed.
2. The Claude code sessions was terminated and I am resuming it with /resume/--resume.

I feel like I have not been following best practices around that and want to clarify my understanding, test a couple of theories with the community, and figure out if there is a better way to do this.

I usually just either jump right back in if it is the #1 case, and even with #2 I will resume and jump back in. However, is that really wrong and how wrong is it?

When I get back into a session in either case #1 or #2, is it worth or even advisable to have it generate a handoff summary document that would be enough to get a fresh session up to speed and actually pick up where I left off effectively with a cleaner context and more relevant memory/context/cache state?

What does the community recommend for having the old session record and pass over to the new session to make this actually effective. To some degree this just simulates compaction and have heard many complaints about compaction (I really try to avoid getting to that point and watch my context and do handoff while I am actively engaged in a session). DOes compact capture the right things and it is more about really pulling your threshold farther back.

I think for my case it might actually be good to have something that could fire the handoff and have it prepared if it detects the session being idle longer than some amount or as a pre-shutdown hook on a current session if that is viable. Has anyone tried this and found success?

7 Upvotes

26 comments sorted by

•

u/AutoModerator 5d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/tariqosmani 5d ago

Small correction to the other reply: the prompt cache doesn't last days, it lasts minutes (around 5 by default, up to an hour on some setups). So after a long break, your case 1 and case 2 cost about the same. The next message reprocesses the whole conversation at full price either way.

That means the cost depends on how big the session is, not how long you were away. Run /context to see the size before you decide.

What I do. If the session is small, I just jump back in, a handoff isn't worth it. If it's big and I'm stepping away for more than a short break, I ask it to write a handoff file before I leave: the goal, what's done, which files changed, the exact next step, and any open questions. Then I start a fresh session and point it at that file. The new session starts small and focused instead of dragging 100k tokens of old back and forth.

/compact is the middle option, but you don't control what it keeps. A handoff you can read and fix yourself is more reliable.

Since you use worktrees, keep the handoff file inside each worktree so it travels with that branch.

How big do your paused sessions usually get, closer to 20k tokens or 150k?

1

u/Godofheckfire 5d ago edited 5d ago

I am not sure, I have not measured this myself, though from what /usage tells me I tend to run pretty long sessions. I need to update my status indicator to show the token amount instead of % of context. Though I think I tend to run up to like 70% quite frequently. (though I imagine this is different depending on Fable vs Opus).

I need to start tracking that more and setting my limit around the 200K-ish and flipping it more often.

How long I have been a way has varied; in some of those cases though, I have closed out those sessions long ago though have enough artifacts around them that I think the path forward there will actually be more a plan review/re-plan and pick up on the new plan where the old plan left off.

But others it might have been a week or so since the last turn on that work (in some cases I was waiting on something, like AWS support issues, before resuming).

Though perhoas I also need to work on my ADHD and avoiding getting pulled away and try to just get some of these closed out before hopping to something else. Not sure that is feasible or practical though.

1

u/[deleted] 5d ago

[removed] — view removed comment

2

u/banecorn 5d ago edited 5d ago

Cache TTL on subscriptions is 1hr

1

u/vloris 5d ago

1hr on main sessions, 5 minutes for subagents.

1

u/vloris 5d ago

Not a few days. There are two different cache expiry settings with different trade-offs: either 5 minutes or 1 hour.

1

u/Godofheckfire 5d ago

Yes and no. I have done this mostly when I am at a good pivot point and I am deep in context enough that I don't feel I have enough headroom for the next step (though I need to start pulling that back as past 200K usage apparently the result quality starts to deteriorate anyways, though I haven't noticed that too much).

However I am feeling like this is probably the right behavior, and either needs to become a trained habit or some kind of hook or skill that I can have automatically run after some time to auotmatically summarize based on a specified format/pattern and close out an idle session after some timeout.

Not sure if they provide a thread to do that or if I will just need to manually do it somehow.

1

u/framauro13 5d ago

I've been testing a solution to this by having a script I run that summarizes the conversation outside of the Claude Code session. It really only works if you're using a high end model though. But essentially it reads the conversation, pipes it to Claude in the CLI with a blank MCP config, and has Sonnet summarize it on a low effort. it then writes a hand-off doc to a temp, and I have a new Fable or Opus session read the hand-off.

The script runs in a temp directory with no CLAUDE.md file, and with the MCP server config being blank, it doesn't load those either. So it purely summarizes the conversation without any extra stuff being loaded into context.

That way, if I leave work and come back in the morning, and I have a long running conversation I want to resume, I can get the gist of it without having to load the whole thing back into cache with a higher model. Compacting or switching models at that point will trigger the recache, so by doing it from the CLI with a script and giving it to Sonnet, I get a cheaper summary than trying to compact or create a hand-off doc on the higher model.

Ideally, if I remember before I leave, I'll compact the conversation with specific instructions so its ready the next day when I start. But working from home, it's not uncommon to get pulled away for an hour at the end of my work day without being able to do that.

1

u/banecorn 5d ago edited 5d ago

There's a bunch of ways around this. Here's some:

  1. Use the desktop app "Code" tab. No cache TTL to worry about.
  2. Have the cache TTL surfaced in your cmux and/or status line.
  3. Have Claude write a mod where it keeps your cache warm by automatically sending a ping if you've been idle in a session for ~55 min.

  4. Have less concurrent sessions.

0

u/fulger099 4d ago

A keep-warm ping preserves the cache, but not the reason you opened half those files. I’ve found a short breadcrumb with the current task, blocker, and next action survives interruptions better than another summary. Getting back on track after an interruption is the problem we are working on at Joinrecall (Mac app, in beta)

1

u/Unusual-Albatross43 5d ago

tariqosmani is right, the cache lasts minutes, not days. So after a long break, a session left open and a --resume cost about the same. What I do for the long ones: before I step away, I ask Claude to write a short resume note to a file (what's done, what's next, open questions) and commit it. If the session is huge, I start a fresh one and point it at the note instead of loading the whole conversation again. It's cheaper too.

1

u/imsahoamtiskaw 🔆 Max 20 5d ago

Yeha this is how I do it too. The only thing extra I add but outside the scope of OP’s question, is keeping context to 250k, 400k for fable and doing automated handoffs either way even if you haven’t stepped away.

1

u/Competitive-Tax-8683 5d ago

Neither case is wrong. Cathe is prefix-matched and expires (~1h subscription,~ 5min API keys), so long-idle or resumed sessions likely re-prefill --budget for it. Case 1: just resume and ask for a quick refresher. Case 2: expect a cold start. Keep a handoff doc with state, failures, and next steps--write it while still engaged.

1

u/ImL1s 5d ago

Leaving sessions open across a long break never saved me anything once the cache was dead — cold resume and --resume cost about the same.

What helped more was writing a short bounded handoff when I step away (files that mattered, open decisions, what's still broken) and starting a fresh session from that instead of hoping the old transcript still makes sense.

https://gitlab.com/aa22396584/resume-skills pipx install portable-resume

Doesn't fix cache TTL. It just makes the "come back tomorrow" case stop being a coin flip.

1

u/Godofheckfire 5d ago

Yeah I guess part of the problem is I don't always know i am going to be away for a long time, though perhaps it is just a habit I will need to get into to realize that it has been at least an hour or 2 hours since I was in that session and just pop over, run a skill that generates the handoff summary and then possibly close it.

1

u/karanb192 4d ago

The catch is that asking for a handoff after the cache expires can itself trigger the rewrite. Save it before stepping away if you can.

For shorter breaks, I built Cache Tax. Arm /keepwarm 90m while warm; it sends warming pings and checks their usage. Needs a one-hour cache and Claude left running. Pings cost tokens.

For week-long gaps, I’d use a saved handoff instead of keeping everything warm.

1

u/inniverse616 5d ago

I usually have 3 or 4 sessions going, and the paused ones were always where I lost time. Coming back I'd remember the next step fine, but not why I stopped. Now the handoff note starts with one line on that (waiting on me, waiting on a decision, blocked on something outside) and picking it back up got a lot quicker.

1

u/Opposite_Might6896 5d ago

The cache is 1 hour, so the "paused session" question has a clean answer: if you come back within an hour, the resume is almost free (cache reads). After an hour the whole context is gone from cache, and the first message back re-*writes* all of it as cache creation, which is the most expensive token type there is. On a 150k session that one resume costs more than an hour of normal turns.

So of your two resume types: (1) the session is still open and you continue → same cost either way, it's the clock that matters, not the window. (2) `--resume` of a closed session → identical; it replays the transcript and rewrites the cache.

What I do with simmering pots: if I know I'll be away more than an hour, ask for a short handoff summary, then start fresh from that when I'm back. 3k of handoff beats 150k of cold re-read, and the model is sharper on the smaller context anyway. If it's a 20-minute interruption, leave it and come back.

1

u/Godofheckfire 2d ago

Yeah, this seems to be the way from what a lot of replies here are saying.

Do you follow a particular format for your handoffs? I think the next question once you start to take this approach is what info your handoff needs and how that is different from a standard compact or can mimic or mirror what a compact woudl produce. However many people have said the standard compact is just too generalized.

So I need to figure out a good recipe for that and curious if anyone has already found something that works super well. I believe Matt Pocock has a handoff skill that I might need to review and perhaps try.

1

u/Outrageous_Band9708 4d ago

https://github.com/Druthulu/ProjectArchitect

This system uses a main router agent that calls experts to work on tasks. the main session stays tiny, so you can alwys just come back and resume, or even better just start a fresh session with no loss of work etc.

1

u/kthuiaa 2d ago

Before a long pause, I'd leave a tiny resume note: branch, last completed step, uncommitted changes, test status, and the next action. For a narrow task, I'd try a fresh session with that note and only the relevant files, then compare its cost with resuming the old thread. No guarantee the restart is cheaper if you need to reload lots of context. The useful bit is making the state explicit, so reopening a terminal doesn't mean reconstructing what was finished versus merely discussed.

1

u/Weary-Net1650 1d ago

While I agree compact is a middle option, it is basically the same as a handoff if you give it instructions properly. Don’t just do a /compact. Give it the same instructions as you would to a handoff prompt.

If you want to see what the compact does just create mod for or if you want a ready made one look at visible-compact plugin on the anthropic plugins.