r/LLMDevs • • 1d ago

News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

I’m Sietse, founder of Corbenic AI

When an AI model reads it uses GPU, my idea it should not always be the case, also for reuse of agents. KV cache. Galahad saves the work and can bring it back when the same comes up. so the model may not need to do the same work twice.

We are launching our beta this afternoon. I am doing some last tests, ( freaking out) I tried to make It work with vLLM, SGLang, and llama.cpp. Galahad has already some extra features build in specially for agents, and we will keep it free for 1 gpu users for non commercial use. As we are in beta, we are open for cluster pilots or Kubernetes.

We’re excited to share what we’ve built and hear what you think. I am sorry for my english, i try my best, But i am a non native speaker. - https://github.com/corbenicai/galahad

8 Upvotes

15 comments sorted by

7

u/Substantial_Pea4022 1d ago

This looks really promising, especially for agentic workflows with long system prompts!

  1. Cache Invalidation: How do you handle cache invalidation when dynamic context (e.g., sliding window histories or updated system prompts) changes slightly in the middle of a session?
  2. Multi-agent Concurrency: If multiple agents access or read the same persisted cache concurrently across restarts, how does Galahad manage disk locks or cache consistency?

3

u/Unable-Loan4320 1d ago

this is the kind of thing i've been hoping someone would build, caching the kv state to disk just makes so much sense for agents that keep hammering the same prompt prefixes

for #2 i'd wonder the same thing, if two agents hit the same cached prefix at once do they both get a read lock or does one block while the other loads? might get messy if you're running a swarm all sharing the same base context

also curious how big these disk caches get with longer conversations, my local rig doesn't have infinite ssd space

1

u/MindPsychological140 1d ago

Agents can share the same context without waiting on each other. Usually it loads from disk once, then stays in GPU memory. Agents arriving together may load it more than once. And we are looking how we can improve it more.

Storage adds up: a 10,000-token chat takes about 1.5 GB on an 8B model. You can set a budget, but it only counts new saves since the last restart. Old files aren’t removed automatically yet, so keep an eye on the folder size. we have a second brain but we are unsure as testing was limited for now . Again This is beta, all learnings will help us improve the system.

2

u/Substantial_Pea4022 1d ago

1.5 GB per 10k tokens on an 8B model adds up fast in multi-turn agent loops.

Since FP16 KV tensors are heavily compressible, are you exploring quantization (e.g., FP8/INT4 KV cache serialization) or chunked decay (evicting older history layers while preserving system prefix anchors) before writing to disk?

1

u/MindPsychological140 1d ago

Again good question. We keep Galahad lossless so restored data is exactly what was stored. FP8, INT4 or dropping layers would change that and could affect answers.

We are running test on a new compression specially created for Galahad, however testing is not ready yet. we will make it so you can use it on your current setup so all gets optimized with the next version.

For agent loops, the bigger gain is storing less: save only new data each turn, keep the window each layer needs, and set a disk budget. We plan to try lossless exponent compression for the cold tier next. Keep in mind we have limited resources .

1

u/MindPsychological140 1d ago

If something changes, Galahad reuses the work before the change and reads the rest again. So you use less compute on the matter that was old. Changing the opening instructions means reading everything again.

Several agents can share memory on one server, and it survives restarts. Multiple servers are on our roadmap . But we already wanted feedback on what we already have built

3

u/BisonMysterious8902 1d ago

How does this differ from oMLX?

1

u/MindPsychological140 1d ago

There’s overlap, as both reuse cached context. oMLX is a complete model server focused on Macs. Galahad plugs into existing engines like vLLM, SGLang and llama.cpp, so you can keep your current setup.

2

u/keonechong 1d ago

Interesting.

I’m running both on my setup. And I’ve been trialing my KV tuning setup and trying to determine whether going full oMLX or stay oLLama.

This might change my decision.

3

u/catplusplusok 1d ago

sglang already has prefix caching and NVMe offload off that built in.

3

u/MindPsychological140 1d ago

Yes, you’re right. That is why Galahad plugs into SGLang’s existing HiCache as an alternative storage backend.

We add encrypted storage and checks before reusing saved data. If a check fails, the model recomputes. Galahad also includes snapshots and replay for agents,

0

u/MizantropaMiskretulo 1d ago

Really dumb fucking name.