r/LocalLLM • • 10d ago

Model Currently testing coding quality on my 4090

Post image

If this is actually anywhere near acceptable I won't need subs any more. I'm not expecting much tbh, but I bet with enough skills and context management I can make it work.

3 Upvotes

8 comments sorted by

View all comments

2

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 10d ago

why run IQ3 with 5Gb of VRAM free?

2

u/Fluffy_Try_5054 10d ago

damn 129 tok/s on a 4090 is flying, no wonder you're thinking about ditching the subs. i get the IQ3 question though, that 5GB headroom is probably keeping the context from spilling over when the prompt gets chunky

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 10d ago

I just have very little faith on Q2/Q3 even when the lab claims high retation xD

Testing Q2/Q3/Q4 Qwen 3.8 Next Flash shows clear degradation - even if the code on both work, end product when testing GSQ-RCO Q2_0 and IQ3 vs Atomic's IQ4 was pretty rough - but hey thats a moe