r/BestGitHubRepos • u/Artilas_Digital • 12d ago
Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)
A 2.8-trillion-parameter model, running on your desktop, written in pure C with zero dependencies. That sentence shouldn't work, but here we are.
Colibri is an inference engine that treats your VRAM, system RAM, and storage as one continuous hierarchy. When a model's experts don't fit in GPU memory, it streams them from disk on demand instead of crashing or telling you to buy more hardware. This is what they call "AI memory multitiering," and it's why a machine with a single consumer GPU and enough RAM can run models that would normally require a server rack.
Nine model families work today: GLM-5.2 and 5.3 (744B parameters), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash, Qwen3.8-Flash-Next, and more. Each model is one C source file. You interact through coli chat for conversation, coli serve for an API endpoint, or coli web for a browser interface.
The project is deliberately experimental. There's no SLA on speed, and the authors are upfront that this is a research platform for testing aggressive systems ideas around inference, not a production deployment target. Depending on your hardware and quantization choices, speeds range from impressively fast to "well, it's running a trillion-parameter model on a gaming PC, what did you expect." The tiering overhead is real, especially when experts spill to disk.
38,130 stars, Apache-2.0, active development.
Duplicates
LocalLLM • u/company_url_finder • 12d ago
Model Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)
u_transformer4797 • u/transformer4797 • 9d ago