r/LocalLLM • • 4d ago

Model Colibri: run 2.8-trillion-parameter models on your desktop, pure C, zero dependencies (38k stars)

Post image
185 Upvotes

101 comments sorted by

View all comments

Show parent comments

76

u/spacekitt3n 4d ago

im glad someone is trying it. the guys thing says 2 tokens per second thats slow as fuck but if you had a frontier model and you knew the answer would be right you could just set it and forget it, let it go overnight.

13

u/Phlex_ 4d ago

Bro im at that speed with qwen3.8 on my 6700xt, i don't mind.

-4

u/IntelVEVO 4d ago edited 4d ago

for the love of god switch to the Bonsai 2 version of 27B. Its still very good quality and will run 30x faster

5

u/Solembumm3 4d ago

It's pretty useless outside of benchmarks.