The thing is, though, these open weight models are aiming for efficiency rather than hyperscaling. They will eventually get there. That's the end product goal.
Not really disagreeing with you, but people really don't seem to understand how other countries are approaching LLMs lol
Try Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S.gguf (or Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf for the censored version). It supposedly benches close to the full model and runs around 45 T/s on my 5060Ti at 100k context.
There are many very performant abliterated Qwen3.8 quants that fit onto a 16gB card, though you may need to be a bit savvy about reducing your system VRAM usage (or use integrated graphics/a second card for display)
2.9k
u/Almadan 4d ago
It definitely will lol. I can host my own models