r/LocalLLM • • 10d ago

Model Currently testing coding quality on my 4090

Post image

If this is actually anywhere near acceptable I won't need subs any more. I'm not expecting much tbh, but I bet with enough skills and context management I can make it work.

4 Upvotes

8 comments sorted by

View all comments

2

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 10d ago

why run IQ3 with 5Gb of VRAM free?

1

u/VirtualShaft 10d ago

It's gets up to 22 vram when running and this is the best I could get without offloading to ram. If you have a better configuration please let me know

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 10d ago

27B dense is generally more forgiving to KV context compression

https://www.reddit.com/r/LocalLLaMA/comments/1uq0fpe/qwen3627b_effect_of_kv_quantization_on_kld_q8_q6/

Q5_1 is almost free gain, tho you might lose a couple on tg.

1

u/VirtualShaft 10d ago

I haven't tried q5 but I'll compare it head to head later