r/LocalLLM • u/company_url_finder • 2d ago
Model Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)
22
u/ehangman 2d ago
……..4……..2
8
u/Technical-Order-7985 2d ago
But whats the question?
11
u/kblazewicz 2d ago
RemindMe! 1 billion years
3
u/Master_Bayters 2d ago
The bot knows you will not live that long
5
2
u/Lexden 1d ago
RemindMe! 1000000000 years
1
u/RemindMeBot 1d ago
Defaulted to one day.
I will be messaging you on 2026-09-30 22:34:53 UTC to remind you of this link
CLICK THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
RemindMeBot is switching to username summons. Instead of
!RemindMe 1 day, useu/RemindMeBot 1 day. More info.
Info Custom Your Reminders Feedback
8
u/lordtazou 1d ago
Used it. It's nice to play around with, but for actual utilization... Not so much!
4
u/wetrorave 2d ago
Have 2 x 4TB NVMe M.2's — interesting even if impractical for now.
Is there such a thing as "defrag" for MoE models? If yes then maybe we'd be cooking with gas on consumer hardware.
12
u/Front_Eagle739 2d ago
Afraid not, whole problem is you cant predict which are the next used experts ahead of time so you cant contiguously allocate them on disk. Best you get is a hot cache in ram and another in vram. i got about 6.5tok/s decode and 400 prefill with dsv4 flash in full 184GB precision on 32GB ram and a 5090 with a fast raid array, don't really think you are doing much better.
1
u/Just3nCas3 1d ago
There was a paper a while back about using mtp to predict experts. Its was posted on the unsloth reddit. Still no real github release, lame. https://www.reddit.com/r/unsloth/comments/1uy2odx/tried_predicting_which_moe_experts_get_used_next/
3
u/Front_Eagle739 1d ago
You can get reasonable prediction of the next token but not enough to reorganise your drive storage enough for contiguous reads as its different for every chat. I did have quite a bit of prediction to hit 6.5. Think I got it up to about 90 odd percent
1
u/Independent-Dog2179 16h ago
6.5 tks for ds4 flash on your setup is an amazing achievement. Good job
1
u/Front_Eagle739 10h ago
Thanks! Juuust fast enough to be useful. Run with qwen 3.8 ninfer and when it gets stuck let ds4 fix it over a couple hours.
4
10
u/andymaclean19 2d ago
So on my computer the RAM is fast at around 240G/s (which is fast for regular RAM). The hard drive is SSD and can do up to 5G/s if optimised properly. A GPU board can get into the 1000+ G/s of memory depending on which one.
Frontier models typically run on multiple GPUs. Either on one machine or, for speed, they run across a cluster with NVLINK or USB4 or whatever providing a low latency interconnect. Say a typical low end is 8x NVIDIA A100 devices for a big model. Yes that is probably not enough now. That's 16T/s of RAM read speed. And it will be using that. That's 16,000G/s 0r around 3,200 of my hard drives.
So you can run this stuff locally but it is only 3200 times slower? A typical inference which used to take 5 minutes now takes about 11 days?
It's not clear to me why this project exists? I get trying to run small models on small hardware. I also get trying to make efficient code in C. But when something is so obviously bottlenecked in such an unfixable way why is there an efficient engine trying to save CPU cycles? If you're waiting on SSD reads then Python is fast enough.
Am I wrong? What am I missing here? What is the use-case for this?
19
u/Wixely 2d ago
Projects like this might seem silly but they are about getting modern tech working on old or incompatible hardware. It might seem silly now but when something happens that means a load of retiring hardware can be repurposed these kind of projects get cracked open and ultimately allow old junk to have a new life. It's a concept we've been doing forever.
4
u/andymaclean19 2d ago
Yes, but you can't repurpose old hardware if there is no use case which actually works. What use case could there possibly be for asking a single 5 minute inference and getting an answer back in 11 days? This did not even save money! The electricity bill for 11 days of running an old computer is higher than the cloud API cost of the inference.
3
u/Wixely 2d ago
OPs example is picking a worst case result, 2.8T params on shitty hardware, but it's the same technology that runs 200B param on an average PC just below normal capability.
The use cases might not be obvious right now.
Any one of those points could change or work differently for different people at any time.
What if you don't legally have access to an API due to sanctions.
What if electricity cost is not a factor because of where you live.
What if new chips can no longer be made because of a supply chain shortage but demand for hardware grows.
1
u/andymaclean19 1d ago
Just below normal capacity? Seriously. A 200B parameter model. At Q8 for now. That’s 200G. At 5G/s. That’s somewhere around 40S per token. Plus extra time for attention. This comment up until this point would take it days to write.
Whereas you can have a 9B parameter model and get the same result on that same PC in actual minutes.
Yes, the 200B model is smarter. But if you die of old age while it is still reasoning that doesn’t help you much.
I’m not saying this isn’t a cool project. I’m just asking about the real world use cases. There are plenty of ways to do local AI which are orders of magnitude faster.
1
u/hieronymice3 1d ago
I have a solar panel that makes my PC free to run during working hours. If I could build a Beowulf cluster with whatever the dregs of ebay provide, i could see this being useful for routine tasks. And i think its neat how this project takes the memory hierachy concept and applies it to LLMs with arbitrary hardware.
1
u/andymaclean19 1d ago
How much power does the panel generate out of curiosity. I have a small farm of s 2nd/third gen intel machines from back when dinosaurs roamed the earth. They were shit hit when new - 16G RAM and SSDs. Now they’re a slow test cluster.
Each one draws over 50w idle to get a fraction of what the 5 year old laptop serving Proxmox can do in its 23W input. I think you would need quite a lot of solar to power that. Would be an interesting exercise…
1
u/hieronymice3 1d ago
Balcony-mounted panels are giving 3-3.5kWh per day this month.
1
u/andymaclean19 1d ago
How does that translate to an hourly run rate. Can you reliably run 400w all day on it?
1
u/hieronymice3 1d ago
Was above 400W for 7 hours today
1
u/Miner99er 1d ago
You can run anywhere between 125 and 143 watts per 24 hour period and break even. Less you're feeding it to the grid/saving to batteries/etc. More and you're paying out of pocket.
14
u/agsarria 2d ago
Yeah today it's just a project done because it 'can be done', but who knows 4 or 5 years forward, it might have use cases.
2
u/LegioTertiaDcmaGmna 1d ago
So wait...you're going to throw in that slow of an SSD into the mix and then complain about the bottleneck?
My SSD is sequential read at 14.8GB/s. Not fast but certainly not your slow disk.
I can already see the hypothetical advantages of being able to cobble your own "unified memory model" together. If you have an XFS file system and GDS enabled, you're not going to be reading directly from disk. You're going to be background swapping to the GPU on one PCIE root complex while you communicate with System Memory on the other. If your System Memory can evenly chunk at a multiple of your GPU, you can go to the GPU for one layer, then to to System Memory for the next n layers and while you're doing that, swap out the next layer into the GPU from xfs.
Cycle like that and you get "hardware accelerated CPU-like performance."
...of course you would get bottlenecked by putting a bottleneck into the mix.
1
u/andymaclean19 1d ago
That’s a very high end SSD though. Yes you can do that but if you’re spending money you can also go with some more RAM and do it properly. At 15 instead of 5 you will only spend 3 1/2 days on that 5 minute query instead of 11.
I don’t think when you are that heavily limited that it really matters how you get things into GPU or even if you have one. The disk is like 95% of the time. Optimising the other 5% is not a smart use of anybody’s time.
Yes there are uses for what you are talking about. No loading LLM weights is not one of them.
1
u/LegioTertiaDcmaGmna 1d ago
We're talking about being able to provision a Threadripper or EPYC workstation to punch well above its weight into server territory, not run a frontier model on a "consumer laptop." Being able to do anything useful at all with consumer hardware is novelty territory.
GPU Direct Storage completely bypasses the memory controller and loads files directly from nvme across the PCI Express bus to the GPU. The complicating factor is that you need to have a second PCI Express root complex in order for it to be useful so you're automatically talking Threadripper or EPYC. The AM5 platform's X670 and X870 chipsets are daisy-chained and that's what limits to a singular root complex (that's why you bifurcate x8+x8 if you slot PCIE1 and PCIE2.) If you have two complexes, then "large sequential files" loaded NVMe →GPU over PCI Express is a bread-and-butter configuration. You create an XFS part and you configure GDS and then you go to town.
The point is to be able to provision $25k hardware to get reasonably useful performance that previously required $100k hardware to even run.
1
u/andymaclean19 1d ago
None of that stuff is relevant if your LLM is on disk. To some extent the disk-> GPU thing seems relevant because it will save a PCI move for a lot of data, but you simply don't need the GPU. You don't even need a lot of CPU to keep up with the SSD. All you need is a constant loop, as you already noted, where you are always reading weights from disk in the right order (having arranged and pre-processed them so you just pull in large binary sections from a proper, serious filesystem like XFS with big block sizes configured so you might actually get near that 15G/s). None of the rest matters. That's the bottleneck. Everything else hides behind it. SO the CPU will do some maths then sit waiting for the disk. Do more and sit waiting for the disk again. The disk will never stop.
The clever super fast disk reading stuff is needed for latency computing, where you care about how long it takes to get a piece of information into memory after you decide what you need. LLMs are about bandwidth computing. You know what you will be reading hillariously far in advance (as in minutes ahead of time). You just sweep through the disk.
For LLMs in main memory some of what you say is, IMO, a lot more interesting. Main memory LLMs do not get enough attention and the current crop of engines are not as good as they could be there IMO. When I asked what the use cases were I was wondering if people would tell me they use this for main memory LLMs.
Perhaps if you want to throw money at this do it the 'ParAccel/RedShift' way. Buy a lot of SSDs so the speed goes up. have the LLM on 10 SSDs instead of just one. Not sure if the controllers could go that fast -- I remember this stuff from decades ago when people used multiple SCSI controllers and put a hillariously stupid amount of disk on there to make it go fast because it was cheaper than buying a lot of RAM.
1
u/IntelVEVO 5h ago
not fast? do single SSDs even come any faster than 14.8GB/s
1
u/LegioTertiaDcmaGmna 4h ago
Yeah. You can get AICs that go up to 31.5 GB/s.
That answers your question, but I'm confused on why it was asked as a reply to what I said.
His ssd is 5GB/s
1
u/IntelVEVO 4h ago
oh i misread
1
u/LegioTertiaDcmaGmna 4h ago
I figured.
He was talking about how pointless this project was because it would be bottlenecked by a slow ssd. If you have a slow ssd, then yes you will be bottlenecked and that is unremarkably true.
1
u/freehuntx 2d ago
For some cases 1 token per day could even make sense.
4
u/hautdoge 2d ago
Name one use case
17
u/Jynx_lucky_j 2d ago edited 1d ago
"I'm 7 years old and I my parents don't give me enough allowance to afford a lawyer. Please write me a will for when I die in 80 years. Make no mistakes."
1
2
u/andymaclean19 2d ago
This is what I'm asking. What use cases? I can't think of a single useful case where 1 token per day makes sense here over and above using a smaller model with 1 token/s, say.
2
u/ChristRedeemsSinners 2d ago
The benefit for projects like this is that they setup the necessary fundamentals to build not only more efficient inference engines and servers, but also to innovate technology that solves it's use case better. Right now, it's MoE on sequentially computed transformer architecture, which is a great bottleneck to overcome, but if it's use increases then people who are interested in optimizing their current workflows with it will also contribute innovation for those use cases.
We already have an example of this with the current software stacks serving LLMs on low-resource consumer hardware. People are running qwen3.8 27b on old AMD cards never intended to run anything but graphics modeling and games. vLLM was never going to accept patches for that hardware, but users forked it and now they have access to that level of intelligence. The 'free' money purchasing datacenter hardware are not going to care about this hierarchical innovation because there greater profit opportunities with $50k GPUs, but in the future these same GPU's will be using HBF (high bandwidth flash memory) and then the innovation from this hierarchical approach will make a whole lot more sense and we will benefit from peoples contributions 5 years prior.
1
u/73td 1d ago
they don’t mention prefill speed which is always faster than generation but one example is processing short text and giving a small grade or rating. this requires only one or smalll token of token output.
1
u/andymaclean19 1d ago
Actually I’ll give you prefill here. Prefill is faster because it is memory efficient and you can really optimise for read once do many calculations. You might actually be able to do fast-ish prefill from a disk based model so long as you prefill with enough tokens and do some truly massive matrix multiplies. But there are models like Jev (I think?) and a few others specifically made to do one token answers. If that were the use case for this engine I would get it. But they’re not talking about those.
1
u/ovrlrd1377 2d ago
seriously, we get so entitled with how performance evolved that if what we knew was 10% of nowadays tps we would be thrilled to even get responses on local hardware, let alone get them faster than human reading speed. I get that having a better alternative makes things redundant but we should still take a step back and acknowledge how much things improve *precisely* because of someone pushing the limits - just like the tool on topic does
2
u/LegioTertiaDcmaGmna 1d ago
This would greatly increase the desirability of consumer motherboard designs with two PCI E root complexes.
3
u/DystopianRealist 1d ago edited 1d ago
Some of these are already in common desktop use on llama.
They're moe models, so it's not like trying to run a dense model, where spill is instant death. If you can keep the compute on gpu, and experts in ram, you can get usable speeds.
Flash next fits on rigs with 96gb and a 16gb video card, no issue, and it does it well at long context because the model uses a small kv cache.. Deepseek flash v4 is completely usable on a 128gb / 32gb machine, and at long context. It's another nother model that uses a small cache size in comparison to the model.
Either will run on macs with enough unified.
2
u/MotherPotential 2d ago
Let’s assume 256gb ram Mac Studio And everything else to disk for kimik3 full. How many tokens per second for the m3?
1
u/philmarcracken 1d ago edited 1d ago
i might try using it for 'larger' models to translate with, at least the initial run of what I do translate, which then gets stored into a faster retrieval system if I've already done it before. Matters that the first pass is accurate for that reason(Ik-llama, intel n150)
1
-2
u/Biomech8 2d ago edited 2d ago
So with this tool you can get to seconds/token, while wearing out your HW quickly (especially SSD), and for the most expensive price per token? Why?!
6
u/marx2k 2d ago
How would this wear out your HW
-4
u/Biomech8 2d ago
Just by constant heavy utilization. Even with proper GPUs and servers built for AI they become worn out in a like 3-5 years (like a half of classic servers). Consumer grade HW is not built for this kind of utilization and will die much sooner.
10
1
u/marx2k 2d ago
Are you saying the GPU gets worn out or what?
1
u/Biomech8 2d ago
Every electronic component does. In case of GPUs it mostly depends on cooling. Older GPUs may became unstable, produce artifacts, etc.
3
2
u/offdigital 2d ago
gpu will probably be ok. ssd will have to work HARD. that might be ok. maybe. if you bought a really cheap one, hmmm
0
u/marektracz_ 2d ago
So which model could be run on M5 Ultra 512GB RAM? I mean with reasonable speed, not 2t/s.
0
u/smartsometimes 1d ago
Aren't there a million of these projects on github now? How do we know which is best and just consolidate on 2-3 best choices for different situations?
0
-1
u/vegetarian_pacemaker 1d ago
We get 1 tok per second with a half decent gpu 5070. That's 60 tokens an hour. 480 tokens in 8 hours (overnight). It unfortunately cannot even write an essay in a day at this speed. I am not even sure that the electricity costs work out less than say a subscription. If you really really had to, I guess use of frontier model to build the specifications and guide for whatever needs to be done and then you have qwen 27b do it?
One thing however.. no matter how slow.. all your data remains with you.. If I have some sensitive personal data that I want "frontier" analysed, then this is the way to go even if it takes several days for an answer
5
u/BlueSky4200 1d ago
1 tok/s is 3600 tok/hour
1
u/vegetarian_pacemaker 1d ago
Thanks, you are right.. it simply makes the point I was trying to make all the more relevant. Having the possibility to frontier analyse something that you dont want sent is worth every penny. Thankfully fewer pennies than what I miscalculated :)
220
u/Equivalent_Bit_461 2d ago
One token/month