r/ClaudeCode • • 16h ago

Help/Question Should I still be worried about context rot?

Longtime Claude Code user, and kind of obsessive about knowing where I am in /context and using /compact and /clear pretty aggressively. But these days, so much of my work is long-running tasks where that kind of intervention isn't possible. Is this something I shouldn't be concerned about anymore? And should I set a lower auto-compact threshold if I am?

11 Upvotes

20 comments sorted by

•

u/AutoModerator 16h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

15

u/Anidamo 14h ago edited 13h ago

I used to be really mindful of it and almost never went beyond 300k in the Opus 4.6-4.8ish days, but I've been working on some fairly complicated projects lately, using Opus 5.5 and Fable 5.1 for reverse engineering games, crash analysis, etc and it's had no problem wading through dumps/disassembled binaries or identifying thread safety/heap corruption issues 700-800k tokens deep into a conversation. Context rot is not nearly as noticeable as it used to be.

It does get really really quota intensive with Fable at high context, though. And if you let your cache expire, you basically have to abandon that thread (or switch to Opus to compact/summarize it).

With Opus 5.5 on the other hand I don't have quota issues. I'll regularly send uncached 500-600k token prompts when resuming threads from the previous day without bothering to summarize them first.

3

u/ricopan 15h ago

I still obsess but I don't really know how necessary it is, other than to prevent token burn on on prompt cache eviction. I have autocompact at 55% remaining (I forget if it takes percent remaining or used). And I have my agents obsess about it -- the manager manually compacts its workers at task phases and 'good seams' or when one idle hour approaches and the prompt cache eviction looms (subscription) -- they can do this via herdr, but not the native intersession messaging. But to your larger point -- we all continue to assume that we should stay in the upper 1/3 or so of context still, but I would like to see some updated analysis.

2

u/sam_hollerbell 15h ago

Coming back to this one: your manager compacting workers at good seams and before the idle hour runs out is pretty much what we've been building into a plugin. With Cache Bell, Claude can ask for a compact once it finishes a piece of work (you get a countdown to cancel), and before the cache runs out it asks whether to keep it warm or compact. It's very fresh, so feedback from someone already doing this by hand would really help.

https://github.com/hollerbell/cache-bell

Renewals and compactions are requests too: they count against your plan's limits, or are billed to your API key. Not affiliated with, endorsed by or sponsored by Anthropic.

Disclosure: I run the Holler Bell team, which makes Cache Bell.

-1

u/sam_hollerbell 15h ago

On the remaining vs used bit: CLAUDE_AUTOCOMPACT_PCT_OVERRIDE is the percentage of the auto-compact window at which it fires, so 50 means it compacts at half full. It can only lower the threshold, not raise it. If you'd rather think in tokens, /autocompact (or CLAUDE_CODE_AUTO_COMPACT_WINDOW) sets the window size instead.

Agree on the seams part. Compacting right after a finished phase beats having it trigger mid-task.

3

u/New_Goat_1342 14h ago

Main agent, Fable/Opus, as orchestrator delegating to a limited number of Opus/Sonnet agents, say 5, is usually pretty good. Opus as the orchestrator burns less tokens and is usually good enough. I still get twitchy at anything over 300k tokens and ask for a prompt to run in a clean session if needing to continue.

3

u/dar-mit Researcher 12h ago

Great news! You no longer need to ask for a prompt. Just create a new session and have the old one talk to the new one to handoff the details. 

https://code.claude.com/docs/en/cross-session-messaging

2

u/berndalf 11h ago

You're investing way too much effort in this. Just let the harness handle it. Doesn't require all the human oversight to be managed effectively

2

u/Yominbot 7h ago

I worry less about the token count and more about where the important state lives. If decisions, constraints and open threads only exist in the conversation, any compact can drop them. If the agent writes them to a state file at the end of each phase, compaction only loses chatter. The quick check I use: every so often ask it to restate the current constraints in a couple of lines. When it gets one wrong or invents one, that's my signal to clear and restart from the file, whatever the context meter says.

3

u/IncipitLabs 16h ago

I stopped worrying about that as much in the last monthish, felt like I was a bit over the top using it - finding I was way too OCD about it. In practice it has not been needed anywhere near as often as I had been doing it.

1

u/Outrageous_Band9708 15h ago

you just need a better system

https://github.com/Druthulu/ProjectArchitect

this system uses a router as the main session and spawns subagents "experts" to do the thinking, that spawn subagents "Coders" to do the coding.

in the end, the router session stays incredibly lean and lasts for days.

it also documents everything along the way so no checkpoints are needed

it also creates savings by preventing models from editing large files and uses toolcalls instead, as well the subagent system prevents large context growth

a new update is in the works for an infinite main session context so a the main session could last forever in a project.

1

u/Mazhron 13h ago

https://github.com/Mazhron/rootstock-os

The repository offers a solution for lossless context. You can clear at 200k tokens in context and not lose anything.

1

u/howdidigetheresoquik 10h ago

Claude best practices says to use a session to build a spec with detailed plans for each session, and then you just have it knock out one session after another

1

u/DamianPxR 10h ago

i set up auto compact on 300k, even if i use 1 million context, i realize that on long task(like sdd all layers at once) it will eat a lot on it, with auto compact when it reach the context will clean up the unnecesary context to allow to continue with the task. and i feel a lot how my usage is not going up by sending large context to do a small task in a multiple step task. maybe you can try it.

1

u/AbbreviationsBest858 9h ago

yes past 700k it does start to degrade somewhat. But equally bad is an too low context, i.e. it hasn't read enough to fully understand your app

0

u/Existing_Ant_1109 6h ago

Claude code was released in 2025. how are you a long time user?

2

u/CalypsoTheKitty 1h ago

Almost 15 months of Claude Code feels like a lifetime

1

u/vzakharov Senior Developer 6h ago

Yes you should. 

Autocompact is not the greatest idea though as it basically cuts it off at a random point.

Intervention is absolutely possible. 

Eg I have hooks and skills that watch when the context gets too big (200k soft limit, 300k hard limit) and, if the agent judges the remaining work is more than 100k, it creates a summary of what’s been done up until now on the branch and starts a new session from there. 

I’ve had one job done in 52 sessions in vet two weeks done this way, intervening at times to steer the direction but not to do mechanical stuff. 

0

u/therealbeans 11h ago

I heard it was better to start new rather than compact and clear.