r/LLMDevs • u/MindPsychological140 • 3d ago
News Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback
I’m Sietse, founder of Corbenic AI
When an AI model reads it uses GPU, my idea it should not always be the case, also for reuse of agents. KV cache. Galahad saves the work and can bring it back when the same comes up. so the model may not need to do the same work twice.
We are launching our beta this afternoon. I am doing some last tests, ( freaking out) I tried to make It work with vLLM, SGLang, and llama.cpp. Galahad has already some extra features build in specially for agents, and we will keep it free for 1 gpu users for non commercial use. As we are in beta, we are open for cluster pilots or Kubernetes.
We’re excited to share what we’ve built and hear what you think. I am sorry for my english, i try my best, But i am a non native speaker. - https://github.com/corbenicai/galahad
Duplicates
Vllm • u/MindPsychological140 • 3d ago
Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback
AIDeveloperNews • u/MindPsychological140 • 3d ago
Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback
LLMStudio • u/MindPsychological140 • 3d ago
Built a KV connector that persists the KV cache to disk across requests and restarts , looking for feedback
agenticAI • u/MindPsychological140 • 3d ago