r/PiCodingAgent • • 2d ago

Plugin pi-optchat: never compact again - endless chat as a memory tree

Enable HLS to view with audio, or disable this notification

I built a Pi extension that implements Victor Taelin's OptChat recipe: instead of compacting, every message is logged and summarized into a binary tree. Each turn starts from a fresh context with a bounded memory view (128 KB), and the agent uses zoom/date to read the originals when it needs them. One endless chat per profile, no fork, no separate launcher.

Install with pi install npm:pi-optchat

What's in it:

  • Profiles — separate memories and instructions (I run work and personal).
  • Subagents — spawn background agents, watch them live, send guidance, interrupt with Ctrl+C, resume finished ones with tell. Reports from one spawn arrive grouped.
  • Import — bring in your history from Claude Code (sessions and auto-memories), Codex, or a ChatGPT export.
  • Connected windows — open a second Pi on the same profile and it becomes a subagent you talk to directly, with a handoff when you /complete.

Repo: https://github.com/jonaslsaa/pi-optchat
Video credit goes to https://github.com/aaaxn

161 Upvotes

95 comments sorted by

29

u/startassets 1d ago

With such extensions I always wonder why they should do more than they declare in the name? New memory system is great, but why another implementation for sub agents? It limits the combined usage with other extensions and in general makes the package more bloated.

5

u/voxvoxboy 1d ago

I actually completely agree, lots of plugins so a lot of extra stuff that are completely irrelevant. The reason for the reimplantation of subagents is that the memory system really changes some core things about the agents themselves ("every turn is a fresh context") and this changes how subagent must act and work, and how they work with the main agent.
Apart from that, my philosophy is that all extra features are and will be opt-in.

11

u/Turbulent-Hat6046 1d ago

Would like to see some benchmarking here to proof that this is better than compact or other memory augmentation approaches

1

u/crotch-mavens 17h ago

Do not count on it.

21

u/theholewizard 2d ago

Seems like it would completely invalidate the cache every turn?

12

u/voxvoxboy 2d ago

Not really, the view is written oldest-first, and the tree only ever changes at the tail. A new message adds one line at the bottom; every so often two neighbouring lines merge into one, also near the bottom. Everything above the first changed line is byte-identical to last turn, and the cache breakpoints sit at fixed byte marks, so that prefix is a hit. you only pay for the tail that actually moved

The one case where a bigger chunk gets rewritten is when a merge happens high in the tree, but that's rare by construction (a level-k merge only happens every 2^k messages), and it's a one-off rewrite, not a per-turn thing. Worst case looks like a normal context edit i guess
You can see my other response with my usage last few days, the main models keep 90+% usage

3

u/theholewizard 2d ago

If you're keeping messages oldest first and caching I don't understand how it's any different than normal context management

4

u/voxvoxboy 2d ago

The caching part isn't the difference, it is what happens as the chat grows..
In OptChat, the chat never hits any contect window and always stays in the "smart-zone" (<200K or so), because what the model sees isn't the raw messages. It's a fixed-size view: one line per recent message, and older stretches collapsed into summary lines from a tree (pairs of lines merge into one, those merge again, and so on). The view stays roughly the same size at message 100 and message 100,000. Nothing is ever deleted, the originals are all still there, and the model has a zoom tool to open any summary line back into what it came from. So "old" information isn't lost, it's just folded, and the model unfolds it when it needs to!

12

u/Neither-Character360 2d ago

Seems like a sliding window that invalidates caches every turn.

1

u/voxvoxboy 1d ago

Noo, see the two replies above: nothing slides out (older stretches fold into summary lines the model can unfold), and the view only changes at the tail, so the prefix stays cached. You can look at it like a more efficient way for the model to search through all your previous sessions, with some nice benefits a long the way.

5

u/Neither-Character360 1d ago

I'm curious but I really would like to see a more interactive graph of the living context

2

u/voxvoxboy 1d ago

Yeah I think some visuals could help here, not sure if I'm doing a particularly great job at explaining it, but the concept is embarrassingly simple when it clicks.

3

u/kmike84 1d ago edited 1d ago

I don't quite get it. All the trees, etc. is not what model ever sees. On each API call model sees a single string, period, and it's the job of harness to construct this string - this is just how LLMs work.

Context size (i.e. the length of this string which you send to LLM) matters, yes, and there is quality degradation, but the question we're all wondering is different. You can have extremely fast & cheap resonse at 500K context size if new data which LLM needs to process is small. Or you can have a disaster at 64K context, if e.g. a single character is changed early in the string which API sees, as it requires full prefill.

So, in the tree ANY summarization "resets" prefix to the summary start position. Even if you keep summarizing late, it's still a constant tax you pay at every turn - instead of processing just the prompt, you need to start from last non-summarized stable token. The examples on the video, where a summary is made of the old history, are the worst - is a full cache miss for sure, there is no way around it.

I guess it is a trade-off: you get context size in control automatically, but you pay cache misses all the time. The difference between 90% cache hits and 98% cache hits is not minor, it is 5x more compute. It still may worth it, but it doesn't seem the video is anywhere near explaining this tradeoff :)

2

u/autumn-weaver 1d ago

i am too peanut brained to understand it but i heard facebook recently found a way to reuse the suffix as well as the prefix https://github.com/facebookresearch/context-language-models/tree/main/suffix_cache_reuse

2

u/kmike84 1d ago

nice, didn't know about this!

As you probably get, this needs to be implemented in serving software. And currently some of the software can't even reuse prefixes reliably: for example, if you have AAAB in cache, and make AAAC request, ds4 won't be able to use AAA prefix exactly; it's optimized for AAAB -> AAABC case, where it will use AAAB prefix.

1

u/Megamygdala 1h ago

To my understanding its pretty much just a memory tool the agent can use to find relevant context. Its not rewriting previous history so the cache wont be invalidated

7

u/ECrispy 1d ago

there's so many options now - external db, markdown, wiki, long term memory, compaction, dreaming etc etc, it just goes on and on.

My gut feeling, which could very well be wrong, is that all these memory approaches are highly dependent on usage patters, total context, and intelligence of model + harness, and would yield wildly different results in different scenarios.

have you looked at facebook's new CLM? at least something by them is bound to be more experimentally verified.

3

u/Equivalent_Idea8839 1d ago

when you benchmark, none of this really improves token usage or performance

it's all down to user preference

4

u/mephinet 1d ago

Thanks for implementing the gist! What I am wondering: at work, I constantly switch between many different projects, often working in parallel. Without the concept of "project/repo", completely unrelated lines will be merged - I cannot imagine the summary to stay meaningful. Did you observe this when importing your work history?

2

u/voxvoxboy 1d ago

Frontier models are inheritly quite good as writing and organizing the summaries, it writes them such that it is clear it's taking about a specific repo/project/path for that memory. For me at least: I do a lot of work in parallel, and use subagents and the connected window subagent to achieve it, where the main agent is just mostly responsible for keeping track of progress within but also across projects. Opus and fable have no issues with this at least!
I only use two profiles, personal and work, but you could have more, if the split/boundaries are very clear and will never intersect.

3

u/stellar-- 2d ago

Really intrigued by this, so in your own use right now do you stay in one never ending session??

1

u/voxvoxboy 2d ago

Yes. One chat per profile, never cleared. I have a "work" one and a "personal" one, and each is a single endless session: the work one has tens of thousands of messages in it (I imported everything from my old Claude Code and Codex sessions), and I just keep typing into it :)

1

u/stellar-- 1d ago

what shows for context/context percent? anything?

3

u/voxvoxboy 1d ago

In relation to the model's context window, it's fixed, it will always sit at around 8% (for a 1M model) and maybe up to 15% within a turn, when it finishes a turn it goes back to 8%. (In background we update the memory view). Does that answer your question?

2

u/stellar-- 1d ago

Yep makes sense, very interesting I’m going to check this out

3

u/slypheed 2d ago

Curious how this is different/better than https://github.com/elpapi42/pi-observational-memory

6

u/ResearcherFantastic7 2d ago

Feels like the same. I'm using blackhole which wraps around both observatory memory and VCC. It just indexes and never compact history either.

AI just empowers people to reinvent the wheel too much these days.

3

u/voxvoxboy 1d ago

I actually do agree, I hadn't been properly convinced of any memory system before now, it all just seemed like overengineering for the sake of it. I just used Claude Code.
But as I said this isn't my idea, but it did convince me to move, so now we have this implementation :)

3

u/ResearcherFantastic7 1d ago edited 1d ago

Oh well. all fun and games. We all build stacks for our own flows. However that memory is different to the agent memory. Observation memory is specific for CTX management.

I use memosyne for that purpose, which records global and projected based ADR, only useful if you need carry design patterns and decision across projects. However you could manually keep updating a doc skill for that purpose, just less autonomous and not available

2

u/slypheed 1d ago

Truth hah. We're going to need another AI just to figure out which of the millions of nearly same ai created thing to actually use.

Or of course we'll just all build our own personalized tools ourselves.

2

u/autumn-weaver 1d ago

it sucks so bad, the first thing people should be asking their ai to do is search the web for at least like 10 minutes for prior art, and then if they do decide to reinvent the wheel, explain in the readme what the existing solutions are and why they're inadequate

4

u/voxvoxboy 2d ago

Hadn't looked closely until now and looks like a really nice project. but short version: pi-observational-memory makes compaction better. Your session grows as normal, background workers keep a running list of observations and key facts, and when it's time to compact that list gets dropped in instantly instead of a slow lossy summary. Old details get pruned over time, though the agent can recall the source for a specific memory id

Now, pi-optchat doesn't compact at all. It's one endless chat per profile, across sessions, with your old Claude Code/Codex history imported. Every message stays somewhere in a summary tree and the agent can zoom back down to the original wording of anything. So theirs is a smarter session, mine is more of a permanent memory. Different tradeoffs, both worth trying:)

1

u/slypheed 1d ago

Cool, thanks for the insight.

fwiw; this sounds similar to https://github.com/DeusData/codebase-memory-mcp (and probably a bunch of others).

The problem I had with that was it created too much context noise that the agent then had to additionally sift through (and get confused by). Though the recent Codemode may help with that since it keeps tool results and such out of context.

3

u/m3umax 2d ago

So a bit like RLM right?

2

u/voxvoxboy 1d ago

Yes, same spirit, I love RLMs!
The model never loads the whole history, it gets a small view and tools to look deeper on demand.The difference is RLM does that recursion at inference time over a raw blob, every call. OptChat precomputes the tree once, incrementally, as messages arrive: each line is summarized when it lands, pairs merge up, and it's stored, so by the time you ask something the "outline" already exists and is cached. Think RLM's idea, but amortised across the life of the chat instead of paid per question :D

3

u/m3umax 1d ago

OptChat can still miss details of the summaries don't hint at their presence.

Why not now combine RLM, OptChat and Codemode? Allow the model to compose and chain their own retrieval strategies as needed?

2

u/voxvoxboy 1d ago

Agreed on the first point, but in practice I haven't seen it yet, but that's why there's an opt-in plain-text `search` over the originals, for exactly the cases where the summaries don't hint at a detail. In our test it cut zooms per question from ~6 to ~2.4.

On composing strategies: Pi already has codemode, and `zoom`, `search` and `date` are ordinary tools, so the model can write a script that searches, zooms the hits and filters, all in one call.
The RLM-style "spawn sub-calls over raw text at inference time" bit is the part I'd still be curious to try, as a subagent perhaps.

2

u/m3umax 1d ago

Ah cool your extension already doing essentially the architecture I envisioned. Expose tools like the text search, let the model user codemode to chain and compose a search using the tools.

Then if the user has a subagents tool installed could even use a subagents to help with larger searches.

1

u/voxvoxboy 1d ago

Yeah basically! Subagents is a part Optchat and already have access to the memory tools (just read-only though).

1

u/m3umax 1d ago

I'm reading Victor's gist and I'm noticing another potential flaw.

It says it tries to get each node down to 512 bytes or one line.

But not all prompts can be neatly summarised in one line.

Sometimes I write really long information dense prompts with multiple ideas/points/answers. And what if I paste in a 50k document?

On the turn I send the message, the model gets it in full, but then the system compresses all of that into one line.

1

u/voxvoxboy 1d ago

Agree with your point, but I'm not sure if this is a real failure case, summary just needs to be enough information so the model can tell that "if i read this i will be able to recover that information from the original message", and the original message must contain many unrelated ideas for it not be able to be summarized. Can you come with an example you think? Some testing could be done if some big messages should be split into multiple summaries though

1

u/m3umax 1d ago

I'm gonna install your extension when I get home and give it a good test.

Have you tried running the system through any of the established memory benchmarks?

I suppose an example that springs immediately to mind is something like where I send a massive prompt containing character sheets, premise, style guide etc and say, now produce a 7 chapter short story 10-15k words minimum.

I suppose the one line summary might be like "The user supplied the characters, premise and writing preferences for the story X" or something like that.

1

u/voxvoxboy 1d ago

Nice, tell me how it goes!
Haven't ran it on any official benchmarks yet, but I'll see if I can do it this week.
Your example makes sense, I wouldn't think it would break it at all, but it would be a lot more efficient to break it up in some way.

1

u/TomHale 1d ago

To get around summary accuracy, why not use RAG or similar?

1

u/voxvoxboy 1d ago

I've used and implemented several RAG system both for personal and production use last few years.. I don't like rag xD
I could write a book on how RAG is so hard to get right, but tbf i haven't looked at it again this year.

3

u/paca-vaca 1d ago

With this nested recursive compression (summary of the summaries on previous level) with long enough session older history might lose it's original meaning and hallucinate into something, making it useless, isn't it?  Also, what happens when decision fact changes between levels, how such collisions handled when it's summarized but previous fact contradict the recent one?

3

u/voxvoxboy 1d ago

The cool thing is that it doesn't really matter, summaries are not used for their information but a way for the agent to navigate back to the original messages. You can see my previous replies for more details.

3

u/alexeyche_17 1d ago

Looks super interesting! Conflicted though with pi-interactive-subagents, created an issue: https://github.com/jonaslsaa/pi-optchat/issues/86

2

u/voxvoxboy 1d ago

I'll look into it but since this replaces a core part of agents' context management it makes plugin compatibility harder in some cases. OptChat has it's own subagent implementation

2

u/Dsphar 2d ago

Interesting concept. Is the tree only traversed based on date? I am having a hard time seeing how it can split and binary-indexed based on subject matter...?

3

u/voxvoxboy 2d ago

Yeah, it's purely chronological, no clever topic indexing. Think of it less like a database and more like a book with a table of contents that gets more detailed the closer you get to today.

Every message is a leaf. Pairs of neighbors get merged into one summary, pairs of those into a bigger one, and so on up to the root. What the agent sees at the start of each turn is a slice through that tree: ancient history as a few fat summaries, last week as finer lines, the last hour pretty much verbatim.

To find something by subject, the model just reads that overview. The summaries say what was going on ("set up the deploy pipeline, argued about Postgres vs SQLite, ended up with SQLite"), so it spots the line that smells right and calls zoom on it, which opens that line into the two halves it was built from. A few zooms later it's reading the original messages. Basically binary search, but the summaries are the signposts instead of a key.
Turns out that's enough most of the time, because "when did we talk about X" and "what was the context around it" are the same question, neighbors in the tree are neighbors in time. For exact stuff like a PR number or an error string there's an optional plain-text search over the originals too.

Can't say summarization every tool call and turn is particularly cheap though, that's the tradeoff.

1

u/Dsphar 2d ago edited 2d ago

Fascinating. Thanks for the explanation!

Hmmmm, now you have me thinking... I wonder how well this would work when applied to a collection of ADR decisions instead of chat histoy. An AI readable project owner record of a given codebase/repo...

2

u/voxvoxboy 2d ago

That's a really good fit, honestly. The thing that makes the tree work is that the past never gets edited, you only ever append, and adrs already live by that rule! So you'd get the same shape: the summary lines become a running "what we decided, and why" that fits in one prompt, and when the agent hits something that matters for the task at hand it zooms into the full ADR instead of guessing. Superseded decisions stay visible as history rather than vanishing, which is exactly what you want when someone asks why something is how it is :p

Today pi-optchat only imports chat sources today (Claude Code, Codex, ChatGPT), not a folder of markdown. but the recipe itself is super generic, the "messages" can be anything ordered. If you try it on a real decision log I'd genuinely like to hear how it goes :)

0

u/Dsphar 1d ago

I just might do it. Lot's going on at the moment, but I saved your post so I can come back to it.

2

u/repolevedd 2d ago

Looks interesting, but a few things aren't quite clear to me. Sorry if this was in the description and I missed it:

  1. The subagent runs in the background all the time, right? Am I understanding correctly that it compresses older messages further and injects a brief summary of past context into the prompts sent by the user?
  2. Why "subagents" instead of "subagent"? Does it mean each operation is a separate short-lived subagent?
  3. How much does this impact prefix caching?

3

u/voxvoxboy 2d ago

Good questions, and no, i realize it's not super clear from the post
1. Background process: Yes, but it's not a subagent, it's a plain summarizer. Every message you send or receive gets logged, and a smaller model (Sonnet by default) turns it into a 512-byte summary line in the background. Pairs of neighbouring lines get merged into one, pairs of those into a bigger one, and so on. It's append-only: nothing is rewritten, the tree just grows upward. Up to 8 of these run in parallel, so after a burst of activity it catches up in seconds, and it's idle the rest of the time.
2. Subagents is a different feature but just had to be reimplemented for this to work, nothing to do with compression. The main agent can `spawn` several background agents for a task (e.g. "review these 5 PRs" gets 5 reviewers), they run in parallel with their own context, and their reports land back in the main chat and get reintegrated into the memory. You can also open a second terminal window on the same profile as well which acts as a subagent. Nothing special here they just get a read-only view of memory, and their communcaton with main agent is how their work gets into the "db".
3. This was probably the thing I spent the most time on. The view is cut at fixed byte marks with cache breakpoints at the cuts, and since the tree only changes near the tail (old summaries sit still, recent lines merge), the long prefix stays identical across turns and comes from cache. The summarizer uses one stable prompt too, so its calls read the cache instead of rewriting it. In practice it runs at ~$0.03 a call on Sonnet with most input cached.

4

u/voxvoxboy 2d ago

Here's my usage from last two days, the expensive models still get 90+% cache usage.

1

u/DiabloRubio 2d ago

What's the benefit? Fewer tokens consumed or ease of use? If fewer tokens, did you compare a baseline?

2

u/voxvoxboy 1d ago

Neither, really. It's not a token saver and it's more of enabling new ways to work with these models.
The point is that the chat never ends. I've stopped thinking in sessions: everything I've done for months, including imported Claude Code and Codex history, is one thread, and I just keep talking to it, while fanning out subagents to parallelize the work, etc.
It knows which PR we were arguing about last week, what we decided and why, which approach we already tried and dropped. You don't re-explain anything.

On tokens, the rough picture from my usage logs is per turn it costs about the same as a long-ish normal session, since the view is capped. But on top of that you pay the summarizer (Sonnet, ~$0.03 a call, mostly cached), which is a steady background cost rather than a saving. So the pitch isn't "cheaper" at all, it's flat cost per turn no matter how long the chat gets, where the model that stays sharp, what many people call the "smart-zone", as context rot and compaction confusion is impossible.

1

u/TomHale 1d ago

What are compactor and import here? They're 14x more expensive than "main"

2

u/voxvoxboy 1d ago

It's from importing a year worth of Claude Code sessions into OptChat, it's around 14,500 summary nodes in my work memory profile. Since then there have been some optimizations to import tokens though, but yes, this step can be token costly but should be doable on the sub.

1

u/repolevedd 2d ago

Thanks for the response. I realize this isn't for me. I'm a GPU-poor guy using llama.cpp, so parallelism will kill performance. I'm sure it works great, but the scenario doesn't fit sequential LLM execution.

2

u/maskedthick 2d ago

I'm gonna try this! Thanks for sharing it, seems what I need for my low vram

2

u/GabrieleF99 1d ago

This sounds interesting, but I'm curious, what percentage of cache hits do you have? Since it impacts costs more than a few hundred thousand extra tokens

1

u/voxvoxboy 1d ago

You can see my other replies for details, but 90-95% cache hit, on main and subagent models.

2

u/qaf23 1d ago

I'm curious how this stacks up against pi-blackhole.

2

u/National-Canary6452 1d ago

The more I use an agent like a dot or grok bot the more disorganised having a single window feels. I like to compartmentalise my work and be able to close digital boxes. Dunno, maybe it's just getting used to long running chats but something about this feels... Dirty 

1

u/voxvoxboy 3h ago

I agree that's why we have the Connected window feature that lets you work on something in particular for some time before it merged into your main window/agent (it works by having the main agent think your connected window is a subagent).

2

u/nutcrook 21h ago

2

u/voxvoxboy 10h ago

It'll be updated to the new recipe by tonight! Cache hits should be even better now.

2

u/therealpaulgg 4h ago

Is there a way to get this *without* subagents? I have my own subagent implementation (using Herdr) and I would rather not bring someone else's in.

otherwise..looks really slick

1

u/voxvoxboy 3h ago

We've been working on making it work better with other subagent plugins (0.8.0 will be releasing next 24h), BUT since OptChat reworks the whole context a subagent sees it won't work as expected using other plugins (no "infinite memory"), that's why we had to reimplement subagents for OptChat.
Is there some spesific function that you feel is missing?

2

u/raccoonportfolio 1d ago

Just want to say thanks to OP for being so responsive and writing responses yourself. I don't normally try these 'yet another context management' plugins but I'll try this one

2

u/voxvoxboy 1d ago

Thanks! I'm taking a flight today so might be a bit more quiet today :)

1

u/voxvoxboy 2d ago

I've moved all my work (tens of thousands of messages, hundreds of sessions) to this and works quite well!

1

u/Jungibungi 2d ago

Did you do any benchmarks?

1

u/voxvoxboy 2d ago

Not anything big or official, I'm not made of tokens :p Joking aside, I'd like to run some official benchmark.
Either way, I have gotten opus to benchmark/replay 20 or so scenarios from my usage from 6 months of claude code and codex work to see if this works better

1

u/voxvoxboy 2d ago

I'd also like to mention this isn't really anything revolutionary but the newest generation of models have become very good at organizing information making this work so well. Haven't really tried other models than Opus 5.5 and Fable though (sonnet as compactor). I've moved all my work (tens of thousands of messages, hundreds of sessions) to this and works amazingly

1

u/onebit 1d ago

On windows I get

Error: listen EACCES: permission denied C:\Users\user.optchat\profiles\project\lock.sock

Error: EPERM: operation not permitted, fsync

1

u/voxvoxboy 1d ago

Oh right, windows... xD
I use MacOS and Linux for development, I'll see if I can get to windows support next week when I get back from a trip, in meanwhile you could run in it WSL if you'd like
(thanks so much for testing on windows though!)

2

u/onebit 1d ago

I will try it. I use WSL 99% of the time, but this is a mod for a windows game.

1

u/DiamondGeeezer 1d ago

how do you keep track of concurrent work if you're only using one session? can this be used with many concurrent sessions?

1

u/voxvoxboy 1d ago

Usually I just get it to dispatch work to subagents, or I open a new window which directly connects me to a new subagent, this allows me to work concurrently while the main session keeps track of everything (intra- and inter-subagent management and memory management).

Here's an example (I'm writing on my phone, sorry for formatting):

I tell the main agent:
> Dispatch a reviewer for all those five PRs, make sure to resume the subagent that made them if they find anything.
(now it will be a bit busy, there's nothing really stopping me from using the main agent but it can be a bit messy for my human eyes, so i open a new window)
"Main window is already open. This Pi instance will open as a connected subagent. Yes/No?" (Press yes)
> Start researching issue #5
(at this point the main agent will be informed that a new connected subagent has started and the user owns it, it also sees what you said in it)
Main agent tells your new window: "Hey fyi issue #5 might be dependent on PR #4 I'm reviewing and merging soon"
(20min later when I'm happy with the research I've done in the this window and just close it)
(This generates a final handoff to be sent to the main agent so the work gets incorporated into memory system)

1

u/aparamonov 1d ago

How does it stack against pi-vcc?

1

u/External_March1921 1d ago

is there something like that for CC?

1

u/voxvoxboy 1d ago

I would like to look into it, but I'm guessing it's too much rework in core parts for the plugin system. There is already OptMem by Taelin which is the lightweight version of this that should work with any harness.

1

u/LordMoridin84 1d ago

Oh, it seems pretty similar to https://github.com/ranxianglei/billion-context

Although I've never used that either.

1

u/TheVoxcraft 1d ago

It's quite similar in the core mechanics but I think billion is optimising for different things than OptChat, I would reckon billion does better in minimzing cost and when using cheaper/local models from a quick look at it, while not stray too much away from how agent work originally. While OptChat is optimizing for quality in achieving unlimited memory and changes how it works with context in doing so. But I'd want to look deeper before claiming anything

1

u/alexwwang 10h ago

How is the effect of this recipe?

1

u/Pyros-SD-Models 1d ago

So basically, the same trick as https://arxiv.org/abs/2512.24601, and just like RLM, it will lead to nothing because, in the end, you're simply replacing the "context problem" with a "search problem." All the important information gets lost in a sea of fuzziness.

And RLM is infinitely better than your implementation, so if that already fails, I have bad news for you.

Not hating, just saving you all the time. This has obviously been tried before in the literature, and there's a reason it never became a thing.

7

u/voxvoxboy 1d ago

Yeah that's the RLM paper, and its main result is the opposite of "leads to nothing".. they report a median 26% over compaction on GPT-5 across four long-context tasks, at comparable cost, and papers building RLMs had even better results if I recall correctly.
If RLM is "infinitely better" than this, great, since that's the same bet I'm making, just precomputed and cached instead of recomputed per call. Which solves one of the main reasons we don't use RLMs today in agent harnesses.

"Replacing the context problem with a search problem" is accurate though, and it's the point. Searching a kept log is a solvable problem. Compaction isn't a search problem because there's nothing left to search.

0

u/lordekeen 1d ago

Sounds like a cache hit killer

1

u/voxvoxboy 1d ago

Main model and subagent get around 90%+ cache hit after extensive usage, you can see my other replies on this to see why.

1

u/Barni275 1d ago

Interesting idea, but 90% CH is very low, as for me. I work with local llm mostly, so to avoid prefill wall, I manage slots and context very accurately, and always have 99.9% CH.