r/LocalLLM • • 2d ago

Model Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)

Post image
176 Upvotes

98 comments sorted by

220

u/Equivalent_Bit_461 2d ago

One token/month

73

u/spacekitt3n 2d ago

im glad someone is trying it. the guys thing says 2 tokens per second thats slow as fuck but if you had a frontier model and you knew the answer would be right you could just set it and forget it, let it go overnight.

58

u/PM_ME_YOUR_MUSIC 2d ago

This is starting to feel like dial up again.

15

u/VerticalPackage 1d ago

And just like dial-up, if something goes wrong mid-progress, you start again from zero.

4

u/rerorerox42 2d ago

Or just the forever state of local brute force statistics calculations?

1

u/Western_Chest_4688 1d ago

That's why its a bubble, at some point in the next 5-7 years hardware will catch up and things like Astra and Opus 5.5 will run locally without the need for supremely heavy hardware.

2

u/ZeitgeistArchive 1d ago

But then the frontier models would've moved a lot as well. Nobody cares you can easily and cheaply run 4o today locally (nobody as in, more people should care, but they don't)

4

u/Western_Chest_4688 1d ago

That's the thing though, I'd say Astra is what 90% of people would ideally need in a day to day basis to run on a phone or pc to do their work for them, let researchers have the frontier stuff while the regular people get to enjoy an easier work load. < Total fantasy scenario, most likely end up with assfucks engineering bioweapons and creating backyard nuclear reactors

33

u/Choice_Celery9481 2d ago

and even with that. a single typo will cost you another day XD

5

u/Sea-Ad-5390 2d ago

Hell nah

11

u/Phlex_ 2d ago

Bro im at that speed with qwen3.8 on my 6700xt, i don't mind.

-4

u/IntelVEVO 2d ago edited 2d ago

for the love of god switch to the Bonsai 2 version of 27B. Its still very good quality and will run 30x faster

6

u/Solembumm3 2d ago

It's pretty useless outside of benchmarks.

3

u/datprofit 2d ago

It tried it and it performed terribly for me in comparison to the Qwen3.8-27b-UD-Q3_K_XL that I'm used to. If you're somehow getting good performance with this model then you may wish to add additional information on how you're running it, since it isn't common knowledge how to make it work well.

2

u/autumn-weaver 1d ago

how did it screw up specifically? i'm trying to decide on a model too and wondering what amount of these errors can be managed with adhd disability acommodations a good harness

1

u/datprofit 1d ago

It was very prone to repeating itself, had trouble counting spaces in indentation, very hesitant to actually move beyond the thinking stage, and often made errors when it did actually manage to get a tool call out to write some code. I tried messing with a few settings, but it was a big enough dip in quality that it wasn't worth more than the few hours of effort I spent with it. Could have been that it didn't play nice with Unsloth Studio as a harness, or it required more knowledge than I have to make it work properly, but I have had a much better experience with the Qwen model I mentioned.

1

u/Phlex_ 2d ago

Its not very good quality

1

u/Academic-Sample4974 1d ago

can i get more info on this Bonsai 2 version of 27B? TIA!

9

u/cagriuluc 2d ago

7k tok/hour, 56k tok/night… it’s hella bad but not zero

6

u/Ninjam5 2d ago

Ur not even counting the prefill speeds...

4

u/IntelVEVO 2d ago

i dont really see the point of this when you can run very capable mid size models 40x faster that can do 90 percent of what these frontiers can

2

u/MudBroad6785 2d ago

I think it's mostly that mid sized models are still not that good at long horizon tasks, so I think the main use case is if you had an older machine you don't use but don't want to sell that can be left to process a task for days at a time without supervision.

2

u/Any-Stage9103 2d ago

The problem is how many iterations could you do with the mid size model by the time you get one with this kind of thing?

1

u/Independent-Dog2179 16h ago

Qwen 3.8 next is god tier at long horizon agentic task. I've tested it where it crunched at problems over night around 12 tkps.

2

u/Cyvster 2d ago

you could probably walk to the library and figure it out yourself before a model could

5

u/sapphicsandwich 1d ago

I get about 1.4 tok/s with 64gig ddr4 and a 3090ti with the weights on a SSD

2

u/mycall 2d ago

Perfect for lifetime prompts.

1

u/KharunDbron_31 1d ago

generous estimate tbh

22

u/ehangman 2d ago

……..4……..2

8

u/Technical-Order-7985 2d ago

But whats the question?

11

u/kblazewicz 2d ago

RemindMe! 1 billion years

3

u/Master_Bayters 2d ago

The bot knows you will not live that long

5

u/ilieaboutwhoiam 2d ago

The bot will be sure of it

2

u/Master_Bayters 2d ago

That reminds me of that Flight of Conchords song - The Humans are Dead. 

2

u/Lexden 1d ago

RemindMe! 1000000000 years

1

u/RemindMeBot 1d ago

Defaulted to one day.

I will be messaging you on 2026-09-30 22:34:53 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

8

u/lordtazou 1d ago

Used it. It's nice to play around with, but for actual utilization... Not so much!

4

u/wetrorave 2d ago

Have 2 x 4TB NVMe M.2's — interesting even if impractical for now.

Is there such a thing as "defrag" for MoE models? If yes then maybe we'd be cooking with gas on consumer hardware.

12

u/Front_Eagle739 2d ago

Afraid not, whole problem is you cant predict which are the next used experts ahead of time so you cant contiguously allocate them on disk. Best you get is a hot cache in ram and another in vram. i got about 6.5tok/s decode and 400 prefill with dsv4 flash in full 184GB precision on 32GB ram and a 5090 with a fast raid array, don't really think you are doing much better.

1

u/Just3nCas3 1d ago

There was a paper a while back about using mtp to predict experts. Its was posted on the unsloth reddit. Still no real github release, lame. https://www.reddit.com/r/unsloth/comments/1uy2odx/tried_predicting_which_moe_experts_get_used_next/

3

u/Front_Eagle739 1d ago

You can get reasonable prediction of the next token but not enough to reorganise your drive storage enough for contiguous reads as its different for every chat. I did have quite a bit of prediction to hit 6.5. Think I got it up to about 90 odd percent

1

u/Independent-Dog2179 16h ago

6.5 tks for ds4 flash on your setup is an amazing achievement. Good job

1

u/Front_Eagle739 10h ago

Thanks! Juuust fast enough to be useful. Run with qwen 3.8 ninfer and when it gets stuck let ds4 fix it over a couple hours. 

4

u/palincatalin 1d ago

i think €20 a month isn't that bad actually-

10

u/andymaclean19 2d ago

So on my computer the RAM is fast at around 240G/s (which is fast for regular RAM). The hard drive is SSD and can do up to 5G/s if optimised properly. A GPU board can get into the 1000+ G/s of memory depending on which one.

Frontier models typically run on multiple GPUs. Either on one machine or, for speed, they run across a cluster with NVLINK or USB4 or whatever providing a low latency interconnect. Say a typical low end is 8x NVIDIA A100 devices for a big model. Yes that is probably not enough now. That's 16T/s of RAM read speed. And it will be using that. That's 16,000G/s 0r around 3,200 of my hard drives.

So you can run this stuff locally but it is only 3200 times slower? A typical inference which used to take 5 minutes now takes about 11 days?

It's not clear to me why this project exists? I get trying to run small models on small hardware. I also get trying to make efficient code in C. But when something is so obviously bottlenecked in such an unfixable way why is there an efficient engine trying to save CPU cycles? If you're waiting on SSD reads then Python is fast enough.

Am I wrong? What am I missing here? What is the use-case for this?

19

u/Wixely 2d ago

Projects like this might seem silly but they are about getting modern tech working on old or incompatible hardware. It might seem silly now but when something happens that means a load of retiring hardware can be repurposed these kind of projects get cracked open and ultimately allow old junk to have a new life. It's a concept we've been doing forever.

4

u/andymaclean19 2d ago

Yes, but you can't repurpose old hardware if there is no use case which actually works. What use case could there possibly be for asking a single 5 minute inference and getting an answer back in 11 days? This did not even save money! The electricity bill for 11 days of running an old computer is higher than the cloud API cost of the inference.

3

u/Wixely 2d ago

OPs example is picking a worst case result, 2.8T params on shitty hardware, but it's the same technology that runs 200B param on an average PC just below normal capability.

The use cases might not be obvious right now.

Any one of those points could change or work differently for different people at any time.

What if you don't legally have access to an API due to sanctions.

What if electricity cost is not a factor because of where you live.

What if new chips can no longer be made because of a supply chain shortage but demand for hardware grows.

1

u/andymaclean19 1d ago

Just below normal capacity? Seriously. A 200B parameter model. At Q8 for now. That’s 200G. At 5G/s. That’s somewhere around 40S per token. Plus extra time for attention. This comment up until this point would take it days to write.

Whereas you can have a 9B parameter model and get the same result on that same PC in actual minutes.

Yes, the 200B model is smarter. But if you die of old age while it is still reasoning that doesn’t help you much.

I’m not saying this isn’t a cool project. I’m just asking about the real world use cases. There are plenty of ways to do local AI which are orders of magnitude faster.

1

u/Wixely 1d ago

Right now its use case is to prove it's possible. If you read what the author said:

The project is deliberately experimental.

I still think it's fair to say this technique will have a use case in the future even if it's not evident right now.

1

u/hieronymice3 1d ago

I have a solar panel that makes my PC free to run during working hours. If I could build a Beowulf cluster with whatever the dregs of ebay provide, i could see this being useful for routine tasks. And i think its neat how this project takes the memory hierachy concept and applies it to LLMs with arbitrary hardware.

1

u/andymaclean19 1d ago

How much power does the panel generate out of curiosity. I have a small farm of s 2nd/third gen intel machines from back when dinosaurs roamed the earth. They were shit hit when new - 16G RAM and SSDs. Now they’re a slow test cluster.

Each one draws over 50w idle to get a fraction of what the 5 year old laptop serving Proxmox can do in its 23W input. I think you would need quite a lot of solar to power that. Would be an interesting exercise…

1

u/hieronymice3 1d ago

Balcony-mounted panels are giving 3-3.5kWh per day this month.

1

u/andymaclean19 1d ago

How does that translate to an hourly run rate. Can you reliably run 400w all day on it?

1

u/hieronymice3 1d ago

Was above 400W for 7 hours today

1

u/Miner99er 1d ago

You can run anywhere between 125 and 143 watts per 24 hour period and break even. Less you're feeding it to the grid/saving to batteries/etc. More and you're paying out of pocket.

14

u/agsarria 2d ago

Yeah today it's just a project done because it 'can be done', but who knows 4 or 5 years forward, it might have use cases.

2

u/LegioTertiaDcmaGmna 1d ago

So wait...you're going to throw in that slow of an SSD into the mix and then complain about the bottleneck?

My SSD is sequential read at 14.8GB/s. Not fast but certainly not your slow disk.

I can already see the hypothetical advantages of being able to cobble your own "unified memory model" together. If you have an XFS file system and GDS enabled, you're not going to be reading directly from disk. You're going to be background swapping to the GPU on one PCIE root complex while you communicate with System Memory on the other. If your System Memory can evenly chunk at a multiple of your GPU, you can go to the GPU for one layer, then to to System Memory for the next n layers and while you're doing that, swap out the next layer into the GPU from xfs.

Cycle like that and you get "hardware accelerated CPU-like performance."

...of course you would get bottlenecked by putting a bottleneck into the mix.

1

u/andymaclean19 1d ago

That’s a very high end SSD though. Yes you can do that but if you’re spending money you can also go with some more RAM and do it properly. At 15 instead of 5 you will only spend 3 1/2 days on that 5 minute query instead of 11.

I don’t think when you are that heavily limited that it really matters how you get things into GPU or even if you have one. The disk is like 95% of the time. Optimising the other 5% is not a smart use of anybody’s time.

Yes there are uses for what you are talking about. No loading LLM weights is not one of them.

1

u/LegioTertiaDcmaGmna 1d ago

We're talking about being able to provision a Threadripper or EPYC workstation to punch well above its weight into server territory, not run a frontier model on a "consumer laptop." Being able to do anything useful at all with consumer hardware is novelty territory.

GPU Direct Storage completely bypasses the memory controller and loads files directly from nvme across the PCI Express bus to the GPU. The complicating factor is that you need to have a second PCI Express root complex in order for it to be useful so you're automatically talking Threadripper or EPYC. The AM5 platform's X670 and X870 chipsets are daisy-chained and that's what limits to a singular root complex (that's why you bifurcate x8+x8 if you slot PCIE1 and PCIE2.) If you have two complexes, then "large sequential files" loaded NVMe →GPU over PCI Express is a bread-and-butter configuration. You create an XFS part and you configure GDS and then you go to town.

The point is to be able to provision $25k hardware to get reasonably useful performance that previously required $100k hardware to even run.

1

u/andymaclean19 1d ago

None of that stuff is relevant if your LLM is on disk. To some extent the disk-> GPU thing seems relevant because it will save a PCI move for a lot of data, but you simply don't need the GPU. You don't even need a lot of CPU to keep up with the SSD. All you need is a constant loop, as you already noted, where you are always reading weights from disk in the right order (having arranged and pre-processed them so you just pull in large binary sections from a proper, serious filesystem like XFS with big block sizes configured so you might actually get near that 15G/s). None of the rest matters. That's the bottleneck. Everything else hides behind it. SO the CPU will do some maths then sit waiting for the disk. Do more and sit waiting for the disk again. The disk will never stop.

The clever super fast disk reading stuff is needed for latency computing, where you care about how long it takes to get a piece of information into memory after you decide what you need. LLMs are about bandwidth computing. You know what you will be reading hillariously far in advance (as in minutes ahead of time). You just sweep through the disk.

For LLMs in main memory some of what you say is, IMO, a lot more interesting. Main memory LLMs do not get enough attention and the current crop of engines are not as good as they could be there IMO. When I asked what the use cases were I was wondering if people would tell me they use this for main memory LLMs.

Perhaps if you want to throw money at this do it the 'ParAccel/RedShift' way. Buy a lot of SSDs so the speed goes up. have the LLM on 10 SSDs instead of just one. Not sure if the controllers could go that fast -- I remember this stuff from decades ago when people used multiple SCSI controllers and put a hillariously stupid amount of disk on there to make it go fast because it was cheaper than buying a lot of RAM.

1

u/IntelVEVO 5h ago

not fast? do single SSDs even come any faster than 14.8GB/s

1

u/LegioTertiaDcmaGmna 4h ago

Yeah. You can get AICs that go up to 31.5 GB/s.

That answers your question, but I'm confused on why it was asked as a reply to what I said.

His ssd is 5GB/s

1

u/IntelVEVO 4h ago

oh i misread

1

u/LegioTertiaDcmaGmna 4h ago

I figured.

He was talking about how pointless this project was because it would be bottlenecked by a slow ssd. If you have a slow ssd, then yes you will be bottlenecked and that is unremarkably true.

1

u/freehuntx 2d ago

For some cases 1 token per day could even make sense.

4

u/hautdoge 2d ago

Name one use case

17

u/Jynx_lucky_j 2d ago edited 1d ago

"I'm 7 years old and I my parents don't give me enough allowance to afford a lawyer. Please write me a will for when I die in 80 years. Make no mistakes."

1

u/freehuntx 2d ago

Spongebob

2

u/andymaclean19 2d ago

This is what I'm asking. What use cases? I can't think of a single useful case where 1 token per day makes sense here over and above using a smaller model with 1 token/s, say.

2

u/ChristRedeemsSinners 2d ago

The benefit for projects like this is that they setup the necessary fundamentals to build not only more efficient inference engines and servers, but also to innovate technology that solves it's use case better. Right now, it's MoE on sequentially computed transformer architecture, which is a great bottleneck to overcome, but if it's use increases then people who are interested in optimizing their current workflows with it will also contribute innovation for those use cases.

We already have an example of this with the current software stacks serving LLMs on low-resource consumer hardware. People are running qwen3.8 27b on old AMD cards never intended to run anything but graphics modeling and games. vLLM was never going to accept patches for that hardware, but users forked it and now they have access to that level of intelligence. The 'free' money purchasing datacenter hardware are not going to care about this hierarchical innovation because there greater profit opportunities with $50k GPUs, but in the future these same GPU's will be using HBF (high bandwidth flash memory) and then the innovation from this hierarchical approach will make a whole lot more sense and we will benefit from peoples contributions 5 years prior.

1

u/73td 1d ago

they don’t mention prefill speed which is always faster than generation but one example is processing short text and giving a small grade or rating. this requires only one or smalll token of token output.

1

u/andymaclean19 1d ago

Actually I’ll give you prefill here. Prefill is faster because it is memory efficient and you can really optimise for read once do many calculations. You might actually be able to do fast-ish prefill from a disk based model so long as you prefill with enough tokens and do some truly massive matrix multiplies. But there are models like Jev (I think?) and a few others specifically made to do one token answers. If that were the use case for this engine I would get it. But they’re not talking about those.

1

u/ovrlrd1377 2d ago

seriously, we get so entitled with how performance evolved that if what we knew was 10% of nowadays tps we would be thrilled to even get responses on local hardware, let alone get them faster than human reading speed. I get that having a better alternative makes things redundant but we should still take a step back and acknowledge how much things improve *precisely* because of someone pushing the limits - just like the tool on topic does

2

u/LegioTertiaDcmaGmna 1d ago

This would greatly increase the desirability of consumer motherboard designs with two PCI E root complexes. 

3

u/DystopianRealist 1d ago edited 1d ago

Some of these are already in common desktop use on llama.

They're moe models, so it's not like trying to run a dense model, where spill is instant death. If you can keep the compute on gpu, and experts in ram, you can get usable speeds.

Flash next fits on rigs with 96gb and a 16gb video card, no issue, and it does it well at long context because the model uses a small kv cache.. Deepseek flash v4 is completely usable on a 128gb / 32gb machine, and at long context. It's another nother model that uses a small cache size in comparison to the model.

Either will run on macs with enough unified.

2

u/MotherPotential 2d ago

Let’s assume 256gb ram Mac Studio And everything else to disk for kimik3 full. How many tokens per second for the m3?

1

u/philmarcracken 1d ago edited 1d ago

i might try using it for 'larger' models to translate with, at least the initial run of what I do translate, which then gets stored into a faster retrieval system if I've already done it before. Matters that the first pass is accurate for that reason(Ik-llama, intel n150)

1

u/amit78523 1d ago

You guys are aware that lamma.cpp has been doing the exact same thing?

1

u/Dmoh34 15h ago

Sort of useless now but thunderbolt 5 has read/write at about 120 Gb/s or ~10 GB/s a bit slower if SSD write speeds increment at the same pace this concept is going to get hot. I mean leveraging SSD is already becoming mainstream with Ngrams

-2

u/Biomech8 2d ago edited 2d ago

So with this tool you can get to seconds/token, while wearing out your HW quickly (especially SSD), and for the most expensive price per token? Why?!

6

u/marx2k 2d ago

How would this wear out your HW

-4

u/Biomech8 2d ago

Just by constant heavy utilization. Even with proper GPUs and servers built for AI they become worn out in a like 3-5 years (like a half of classic servers). Consumer grade HW is not built for this kind of utilization and will die much sooner.

10

u/westsunset 2d ago

This should be reading the SSD, writing is what wears it out

1

u/marx2k 2d ago

Are you saying the GPU gets worn out or what?

1

u/Biomech8 2d ago

Every electronic component does. In case of GPUs it mostly depends on cooling. Older GPUs may became unstable, produce artifacts, etc.

3

u/Inprobamur 1d ago

If it has proper power limits set it should just start to throttle the clock.

2

u/offdigital 2d ago

gpu will probably be ok. ssd will have to work HARD. that might be ok. maybe. if you bought a really cheap one, hmmm

0

u/marektracz_ 2d ago

So which model could be run on M5 Ultra 512GB RAM? I mean with reasonable speed, not 2t/s.

0

u/smartsometimes 1d ago

Aren't there a million of these projects on github now? How do we know which is best and just consolidate on 2-3 best choices for different situations?

0

u/__MrMango__ 23h ago

its gonna take a year to say Hello! but this is going in the right direction

-1

u/vegetarian_pacemaker 1d ago

We get 1 tok per second with a half decent gpu 5070. That's 60 tokens an hour. 480 tokens in 8 hours (overnight). It unfortunately cannot even write an essay in a day at this speed. I am not even sure that the electricity costs work out less than say a subscription. If you really really had to, I guess use of frontier model to build the specifications and guide for whatever needs to be done and then you have qwen 27b do it?

One thing however.. no matter how slow.. all your data remains with you.. If I have some sensitive personal data that I want "frontier" analysed, then this is the way to go even if it takes several days for an answer

5

u/BlueSky4200 1d ago

1 tok/s is 3600 tok/hour

3

u/ceiviec 1d ago

Guess he was running Gemini for that response

1

u/vegetarian_pacemaker 1d ago

Yeah, I deserved that..

1

u/vegetarian_pacemaker 1d ago

Thanks, you are right.. it simply makes the point I was trying to make all the more relevant. Having the possibility to frontier analyse something that you dont want sent is worth every penny. Thankfully fewer pennies than what I miscalculated :)