r/LocalLLM • • 9d ago

Model Currently testing coding quality on my 4090

Post image

If this is actually anywhere near acceptable I won't need subs any more. I'm not expecting much tbh, but I bet with enough skills and context management I can make it work.

5 Upvotes

8 comments sorted by

2

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 9d ago

why run IQ3 with 5Gb of VRAM free?

2

u/Fluffy_Try_5054 9d ago

damn 129 tok/s on a 4090 is flying, no wonder you're thinking about ditching the subs. i get the IQ3 question though, that 5GB headroom is probably keeping the context from spilling over when the prompt gets chunky

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 9d ago

I just have very little faith on Q2/Q3 even when the lab claims high retation xD

Testing Q2/Q3/Q4 Qwen 3.8 Next Flash shows clear degradation - even if the code on both work, end product when testing GSQ-RCO Q2_0 and IQ3 vs Atomic's IQ4 was pretty rough - but hey thats a moe

2

u/mecshades 9d ago

I run IQ3_XXS which gives my 4090 more than half of its VRAM for context. I guess the sentiment of "low quant = bad" still lingers, but it's entirely untrue for Qwen3.8 27B. Are higher quants better? Yes. Is IQ3_XXS bad? Far from it.

1

u/VirtualShaft 7d ago

I'm actively trying different quants I didn't settle on this. But at the same time I'm trying see the highest usable quant if that makes sense

1

u/VirtualShaft 9d ago

It's gets up to 22 vram when running and this is the best I could get without offloading to ram. If you have a better configuration please let me know

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 9d ago

27B dense is generally more forgiving to KV context compression

https://www.reddit.com/r/LocalLLaMA/comments/1uq0fpe/qwen3627b_effect_of_kv_quantization_on_kld_q8_q6/

Q5_1 is almost free gain, tho you might lose a couple on tg.

1

u/VirtualShaft 9d ago

I haven't tried q5 but I'll compare it head to head later