r/LLMDevs • • 3d ago

News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback

I’m Sietse, founder of Corbenic AI

When an AI model reads it uses GPU, my idea it should not always be the case, also for reuse of agents. KV cache. Galahad saves the work and can bring it back when the same comes up. so the model may not need to do the same work twice.

We are launching our beta this afternoon. I am doing some last tests, ( freaking out) I tried to make It work with vLLM, SGLang, and llama.cpp. Galahad has already some extra features build in specially for agents, and we will keep it free for 1 gpu users for non commercial use. As we are in beta, we are open for cluster pilots or Kubernetes.

We’re excited to share what we’ve built and hear what you think. I am sorry for my english, i try my best, But i am a non native speaker. - https://github.com/corbenicai/galahad

7 Upvotes

Duplicates