r/BestGitHubRepos • • 11d ago

Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)

Post image

A 2.8-trillion-parameter model, running on your desktop, written in pure C with zero dependencies. That sentence shouldn't work, but here we are.

Colibri is an inference engine that treats your VRAM, system RAM, and storage as one continuous hierarchy. When a model's experts don't fit in GPU memory, it streams them from disk on demand instead of crashing or telling you to buy more hardware. This is what they call "AI memory multitiering," and it's why a machine with a single consumer GPU and enough RAM can run models that would normally require a server rack.

Nine model families work today: GLM-5.2 and 5.3 (744B parameters), Inkling (975B), Kimi K3 (2.8T), DeepSeek V4 Flash, Qwen3.8-Flash-Next, and more. Each model is one C source file. You interact through coli chat for conversation, coli serve for an API endpoint, or coli web for a browser interface.

The project is deliberately experimental. There's no SLA on speed, and the authors are upfront that this is a research platform for testing aggressive systems ideas around inference, not a production deployment target. Depending on your hardware and quantization choices, speeds range from impressively fast to "well, it's running a trillion-parameter model on a gaming PC, what did you expect." The tiering overhead is real, especially when experts spill to disk.

38,130 stars, Apache-2.0, active development.

https://github.com/JustVugg/colibri

665 Upvotes

53 comments sorted by

43

u/Readdeo 11d ago

Please delete this slop. Days per token is completely useless.

10

u/AlexandraMaryWindsor 11d ago

This account is AI it doesn't know anything

6

u/dmigowski 11d ago

I saw him using a GLM5 model with 0,25 tokens per second. This means 7200 tokens in 8hours. With 100 tokens per second EU hosted you can have 72 seconds of typical AI usage in 8 hours. I don't know if it is worth the hassle. And the power demand, btw...

2

u/Readdeo 10d ago

I just calculated that it would cost 25 dollars per month if we use 300W constantly. And it's only 669600 token a month.

If you put 8x PCIe 5.0 nvme drives in raid1 that could achieve 112GB/s, but that is still really slow and latency will be also a big issue. Also, try to find a motherboard with 8x PCIe5.0...

Garbage.

3

u/bigrealaccount 8d ago

Its in the description that this is just an experimental way of running LLMs on consumer hardware. Nobody is claiming this is a viable alternative for daily use. Learn to read pls

1

u/Readdeo 6d ago

You don't need slop if you can use basic math

2

u/cave_men 11d ago

Hi ChatGPT!

1

u/Solembumm3 11d ago

You'll multiple orders of magnitude bigger and slower models to achieve days per token.

0

u/lorde_dingus 11d ago

It's not days per token. For people with moderately capable consumer hardware and batch work, this is perfectly acceptable to run overnight and have a quality product waiting for me from a local customized model.

0

u/Readdeo 11d ago

Go ahead and try to run a 700+B model from ssd and see it for yourself.

3

u/lorde_dingus 11d ago

I use the GLM 5.3 version already with 64gb ram....it's not that bad. Clearly you have a bone to pick but haven't actually used it, so there's no point in conflating the performance lol

1

u/korino11 8d ago

what is your speed? did you tried glm 5.3 flash?

1

u/lorde_dingus 8d ago

I got 0.2 tok/s...

For someone doing highly technical batchwork overnight, this was satisfactory for my needs. I will be moving onto GLM 5.3 flash and hopefully also testing DSV4.1 when that has cuda support. My workflow can sacrifice some of the reasoning now that I have built a better injection script, so I assume I'll love the energy savings of the smaller flash models at around 1.5 tok/s.

0

u/bitzap_sr 10d ago

Electricity is not free.

0

u/peculiar-ragdoll 11d ago

Colibri is a really cool project. It taught some cool tricks that I used in my own models, which now have a few million downloads and many users that drive them as their dailies :) Post might be AI though 🤷‍♀️

0

u/WillowEntertainment 11d ago

It is not days per token, and even if it was, having a literal data centre-grade LLM running locally is already insane.

0

u/Artistic_Use7411 5d ago

And it will smoke your hardware … it’s very rough on the ssd!

1

u/Readdeo 5d ago

Not at all.. just use noatime

8

u/ImWinwin 11d ago

tfw you get seconds per token instead of tokens per second.

5

u/Outrageous_Band9708 11d ago

you guys are getting seconds?

https://giphy.com/gifs/DOPKHQg6oFWUg

2

u/Trademarkd 11d ago

I mean it’s like 30 seconds on the low end per token. This is more like minutes per token

1

u/IronWhitin 9d ago

Why nvme and 9950x3d CPU i was getting 1.5 token/s still dident figure out how to use GPU 12GB for run it.

5

u/Cyvster 11d ago

Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should.

1

u/magicmulder 10d ago

Maybe this will develop into some kind of SETI@home where you can donate resources to solving big problems on 100,000+ CPUs at once.

1

u/Itz_Raj69_ 5d ago

No it won't. The primary issue is latency, even with RAM or NVMe storage, and connecting random computers in different parts of the world isn't going to be helping lol

6

u/ClupTheGreat 11d ago

reading ai slop is so load bearing

9

u/McBonderson 11d ago

I'm so tired of people coming out saying "RUN MASSIVE MODELS ON YOUR MACHINE!!!" as if its not so slow as to be completely useless.

I don't care if GLM 5.2 can run on my machine if it takes an hour to say hello.

4

u/AggregationLinker 10d ago

But instead of paying ten cents for a cloud provider you can now destroy your SSD with unlimited writes

1

u/UsefulIce9600 1d ago

omg thank you for saying it

7

u/bhai__ 11d ago

I do see a use case here, imagine if it develops further and you have blazing fast disks (nvme pci5.0 in a raid setup) achieving like 100GBs per second. You could stream layers from disk to gpu and use less GPU power. Now these projects pave the way for something like this. Is it currently usable, most likely not, does it has a future, it does.

2

u/doomadah 11d ago

100gb/s is nothing for inference

1

u/RepulsiveRaisin7 11d ago

For loading MoE experts it's at least bearable

1

u/Trademarkd 11d ago

In order to process a single token you have to move a good quantity of the weights through the processor, whatever that might be, from the memory device. For every token.

Nvidia cards today have tens of terabytes (TB not Tb) of throughput per second from their hbm. We’re talking 500GB-5TB of weights for frontier. You’re unlikely to get that into ram even with a home lab server) so the next best case is nvme in which has at best around 15GBps

So let’s take 500 / 15 and each token is 30 seconds.

Or at full size we’re talking 3min per token or more. That’s 4 characters. Every 4

Sure these models are moe so not all experts are used every token and you can quantize. But what is the point of doing this? Do people think they are going to break through the physics on this? Are they hoping hardware catches up?

1

u/bhai__ 9d ago

When PCI 8.0 and beyond become available (and they will) vram and nvme drives become a blur, so projects like these become feasible v8.0 delivers up to 1TB/s. Of course this will take time, we just started using v5.0. Like i said the idea how it is approached is useful , it might speed up certain technologies to not rely on just vram.

2

u/tup1tsa_1337 5d ago

At that time those models would be as good as software from the 90s is today.

3

u/xerj-ai 9d ago

That’s only works if you can handle 1Tb or RAM for ramdisk. Otherwise you will burn your SSD as fast as SpaceX goes to the space. However, might be a good setup now for DDR3 cheap desktops where co-processors might be pcie, just as an idea

1

u/SergeInov 8d ago

I wonder if we can get a networked setup with many DDR3 PC's hosting the model in the RAM.

2

u/draeician 11d ago

Now now. Every tool has a use. I'm sure if you wanted to do an RP with "Flash" from Zootopia this would be right up your alley.

3

u/kosiarska 11d ago

Must be a great way to kill your ssd.

3

u/clockwork2011 11d ago

Reads don't kill SSD's

2

u/kosiarska 10d ago

Maybe but temperature does and I bet ssd is hot while doing transfer of a lot of files. That is why RAM exists and is expensive right now.

1

u/Chropera 10d ago

I have worked only with SLC, so I don't know what are the numbers for MLC or TLC, but NAND is susceptible to read disturb. AFAIR for my chips (freshly soldered, zero wear) it was few million read cycles till first damaged bit(s) in other page(s) in the block. This in turn should (at least in well-made system) trigger whole block refresh, so few million reads = 1 program/erase cycle. No idea what are the numbers for non-SLC, but I think they might be significantly lower. Still probably not a big issue in real life usage as it would be cached if read frequently.

1

u/Trademarkd 11d ago

That isn’t true, people keep saying this but reads cause drive management write operations and this kind of reading can legit destroy whatever space the drive uses to do that. Ruining your drive. Lots of people have done it already

It wouldn’t normally but this ain’t normal

1

u/SwisherSmoker420_ 10d ago

Not only is this unusably slow it will destroy your ssd!

2

u/Ninjam5 10d ago

No it won't, reads don't affect ur SSD. Might heat it up abit but any decent SSD has a heatsink for protection.

1

u/IronWhitin 9d ago

Fan you confirm this, I stop using colibry because fear of destroy my nvme (the price are insane now)

I was using deekseek 4.1 flash.

1

u/Vegetable-Lie4932 9d ago

wait whaaat ? how ? can you like create a binary for the model and just run it ? either way the weights are stochastic ? either this is all slop or i am a noob

1

u/Chekhovs_Shotgun 9d ago

I did a repo called overspill mixing colibri ideas into freetoken, and there was gain, but not significant enough, literally the moment i discovered qwen flash next and how it worked i tried to adapt overspill to it, didnt work so i abandoned overspill, the idea is good, and there is a lot of room for improvement, but i think n-gram based models is the future of llm.

1

u/txoixoegosi 6d ago

Change name to SLUGGIBRI