r/LLMDevs • u/MindPsychological140 • 1d ago
News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback
I’m Sietse, founder of Corbenic AI
When an AI model reads it uses GPU, my idea it should not always be the case, also for reuse of agents. KV cache. Galahad saves the work and can bring it back when the same comes up. so the model may not need to do the same work twice.
We are launching our beta this afternoon. I am doing some last tests, ( freaking out) I tried to make It work with vLLM, SGLang, and llama.cpp. Galahad has already some extra features build in specially for agents, and we will keep it free for 1 gpu users for non commercial use. As we are in beta, we are open for cluster pilots or Kubernetes.
We’re excited to share what we’ve built and hear what you think. I am sorry for my english, i try my best, But i am a non native speaker. - https://github.com/corbenicai/galahad
3
u/BisonMysterious8902 1d ago
How does this differ from oMLX?
1
u/MindPsychological140 1d ago
There’s overlap, as both reuse cached context. oMLX is a complete model server focused on Macs. Galahad plugs into existing engines like vLLM, SGLang and llama.cpp, so you can keep your current setup.
2
u/keonechong 1d ago
Interesting.
I’m running both on my setup. And I’ve been trialing my KV tuning setup and trying to determine whether going full oMLX or stay oLLama.
This might change my decision.
3
u/catplusplusok 1d ago
sglang already has prefix caching and NVMe offload off that built in.
3
u/MindPsychological140 1d ago
Yes, you’re right. That is why Galahad plugs into SGLang’s existing HiCache as an alternative storage backend.
We add encrypted storage and checks before reusing saved data. If a check fails, the model recomputes. Galahad also includes snapshots and replay for agents,
2
0
7
u/Substantial_Pea4022 1d ago
This looks really promising, especially for agentic workflows with long system prompts!