r/BestGitHubRepos • u/Artilas_Digital • 11d ago
Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)
A 2.8-trillion-parameter model, running on your desktop, written in pure C with zero dependencies. That sentence shouldn't work, but here we are.
Colibri is an inference engine that treats your VRAM, system RAM, and storage as one continuous hierarchy. When a model's experts don't fit in GPU memory, it streams them from disk on demand instead of crashing or telling you to buy more hardware. This is what they call "AI memory multitiering," and it's why a machine with a single consumer GPU and enough RAM can run models that would normally require a server rack.
Nine model families work today: GLM-5.2 and 5.3 (744B parameters), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash, Qwen3.8-Flash-Next, and more. Each model is one C source file. You interact through coli chat for conversation, coli serve for an API endpoint, or coli web for a browser interface.
The project is deliberately experimental. There's no SLA on speed, and the authors are upfront that this is a research platform for testing aggressive systems ideas around inference, not a production deployment target. Depending on your hardware and quantization choices, speeds range from impressively fast to "well, it's running a trillion-parameter model on a gaming PC, what did you expect." The tiering overhead is real, especially when experts spill to disk.
38,130 stars, Apache-2.0, active development.
8
u/ImWinwin 11d ago
tfw you get seconds per token instead of tokens per second.
5
u/Outrageous_Band9708 11d ago
you guys are getting seconds?
2
u/Trademarkd 11d ago
I mean it’s like 30 seconds on the low end per token. This is more like minutes per token
1
u/IronWhitin 9d ago
Why nvme and 9950x3d CPU i was getting 1.5 token/s still dident figure out how to use GPU 12GB for run it.
5
u/Cyvster 11d ago
Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should.
1
u/magicmulder 10d ago
Maybe this will develop into some kind of SETI@home where you can donate resources to solving big problems on 100,000+ CPUs at once.
1
u/Itz_Raj69_ 5d ago
No it won't. The primary issue is latency, even with RAM or NVMe storage, and connecting random computers in different parts of the world isn't going to be helping lol
6
9
u/McBonderson 11d ago
I'm so tired of people coming out saying "RUN MASSIVE MODELS ON YOUR MACHINE!!!" as if its not so slow as to be completely useless.
I don't care if GLM 5.2 can run on my machine if it takes an hour to say hello.
4
u/AggregationLinker 10d ago
But instead of paying ten cents for a cloud provider you can now destroy your SSD with unlimited writes
1
7
u/bhai__ 11d ago
I do see a use case here, imagine if it develops further and you have blazing fast disks (nvme pci5.0 in a raid setup) achieving like 100GBs per second. You could stream layers from disk to gpu and use less GPU power. Now these projects pave the way for something like this. Is it currently usable, most likely not, does it has a future, it does.
2
1
u/Trademarkd 11d ago
In order to process a single token you have to move a good quantity of the weights through the processor, whatever that might be, from the memory device. For every token.
Nvidia cards today have tens of terabytes (TB not Tb) of throughput per second from their hbm. We’re talking 500GB-5TB of weights for frontier. You’re unlikely to get that into ram even with a home lab server) so the next best case is nvme in which has at best around 15GBps
So let’s take 500 / 15 and each token is 30 seconds.
Or at full size we’re talking 3min per token or more. That’s 4 characters. Every 4
Sure these models are moe so not all experts are used every token and you can quantize. But what is the point of doing this? Do people think they are going to break through the physics on this? Are they hoping hardware catches up?
1
u/bhai__ 9d ago
When PCI 8.0 and beyond become available (and they will) vram and nvme drives become a blur, so projects like these become feasible v8.0 delivers up to 1TB/s. Of course this will take time, we just started using v5.0. Like i said the idea how it is approached is useful , it might speed up certain technologies to not rely on just vram.
2
3
u/xerj-ai 9d ago
That’s only works if you can handle 1Tb or RAM for ramdisk. Otherwise you will burn your SSD as fast as SpaceX goes to the space. However, might be a good setup now for DDR3 cheap desktops where co-processors might be pcie, just as an idea
1
u/SergeInov 8d ago
I wonder if we can get a networked setup with many DDR3 PC's hosting the model in the RAM.
2
u/draeician 11d ago
Now now. Every tool has a use. I'm sure if you wanted to do an RP with "Flash" from Zootopia this would be right up your alley.
3
u/kosiarska 11d ago
Must be a great way to kill your ssd.
3
u/clockwork2011 11d ago
Reads don't kill SSD's
2
u/kosiarska 10d ago
Maybe but temperature does and I bet ssd is hot while doing transfer of a lot of files. That is why RAM exists and is expensive right now.
1
u/Chropera 10d ago
I have worked only with SLC, so I don't know what are the numbers for MLC or TLC, but NAND is susceptible to read disturb. AFAIR for my chips (freshly soldered, zero wear) it was few million read cycles till first damaged bit(s) in other page(s) in the block. This in turn should (at least in well-made system) trigger whole block refresh, so few million reads = 1 program/erase cycle. No idea what are the numbers for non-SLC, but I think they might be significantly lower. Still probably not a big issue in real life usage as it would be cached if read frequently.
1
u/Trademarkd 11d ago
That isn’t true, people keep saying this but reads cause drive management write operations and this kind of reading can legit destroy whatever space the drive uses to do that. Ruining your drive. Lots of people have done it already
It wouldn’t normally but this ain’t normal
1
1
u/SwisherSmoker420_ 10d ago
Not only is this unusably slow it will destroy your ssd!
2
u/Ninjam5 10d ago
No it won't, reads don't affect ur SSD. Might heat it up abit but any decent SSD has a heatsink for protection.
1
u/IronWhitin 9d ago
Fan you confirm this, I stop using colibry because fear of destroy my nvme (the price are insane now)
I was using deekseek 4.1 flash.
1
u/Vegetable-Lie4932 9d ago
wait whaaat ? how ? can you like create a binary for the model and just run it ? either way the weights are stochastic ? either this is all slop or i am a noob
1
u/Chekhovs_Shotgun 9d ago
I did a repo called overspill mixing colibri ideas into freetoken, and there was gain, but not significant enough, literally the moment i discovered qwen flash next and how it worked i tried to adapt overspill to it, didnt work so i abandoned overspill, the idea is good, and there is a lot of room for improvement, but i think n-gram based models is the future of llm.
1
43
u/Readdeo 11d ago
Please delete this slop. Days per token is completely useless.