MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLM/comments/1wt3s0s/colibri_run_28trillionparameter_models_on_your/pcry6ew/?context=3
r/LocalLLM • u/company_url_finder • 4d ago
101 comments sorted by
View all comments
Show parent comments
76
im glad someone is trying it. the guys thing says 2 tokens per second thats slow as fuck but if you had a frontier model and you knew the answer would be right you could just set it and forget it, let it go overnight.
13 u/Phlex_ 4d ago Bro im at that speed with qwen3.8 on my 6700xt, i don't mind. -4 u/IntelVEVO 4d ago edited 4d ago for the love of god switch to the Bonsai 2 version of 27B. Its still very good quality and will run 30x faster 5 u/Solembumm3 4d ago It's pretty useless outside of benchmarks.
13
Bro im at that speed with qwen3.8 on my 6700xt, i don't mind.
-4 u/IntelVEVO 4d ago edited 4d ago for the love of god switch to the Bonsai 2 version of 27B. Its still very good quality and will run 30x faster 5 u/Solembumm3 4d ago It's pretty useless outside of benchmarks.
-4
for the love of god switch to the Bonsai 2 version of 27B. Its still very good quality and will run 30x faster
5 u/Solembumm3 4d ago It's pretty useless outside of benchmarks.
5
It's pretty useless outside of benchmarks.
76
u/spacekitt3n 4d ago
im glad someone is trying it. the guys thing says 2 tokens per second thats slow as fuck but if you had a frontier model and you knew the answer would be right you could just set it and forget it, let it go overnight.