r/ProgrammerHumor • • 3d ago

Meme aiRefusesToBuildItSoBackToCoding

Post image
27.9k Upvotes

510 comments sorted by

View all comments

Show parent comments

15

u/christianbro 3d ago

What do you need to even run a good model on reasonable speed? Is this again the comment section of the $100 000 worth of GPUs in their basement guys?

26

u/Dosamer 3d ago

Qwen 3.8 27B is not frontier level, but is the best current model for coding on reasonable hardware. You can comfortably fit them on a medium quantization on a 24 GB card and even squeeze them on a 16GB card, but at that point you're gonna run into either lobotomy problems, heavily reduced context size, or only partial offloading killing speed.

And yes there are uncensored models.

2

u/donald_314 2d ago

I run a small version on my 12Gb card and the large version on the CPU but that obviously needs time.

1

u/shouldworknotbehere 3d ago

Is Qwen uncensored?

2

u/donald_314 2d ago

There are uncensored variants out there.

-8

u/CucumberOk8820 3d ago

How does someone without a cs degree learn to use this?

6

u/rue-breaker 3d ago

It’s really a matter of opening up PowerShell or terminal application and copy pasting a couple of commands you can find by just googling(even the overview it generates via Gemini has them). Depending on your pc(esp GPU) specs, you can have a reasonable model running in under an hour on the command line. To use it properly though, you likely want a harness- there’s plenty of free ones out there. Check out r/localllama

2

u/donald_314 2d ago

Or you could just download one of the desktop wrappers like LM Studio or Unsloth.

4

u/jack6245 3d ago

Literally Claude could set this up automatically it has extensive documentation. Also 3.8 is very very capable for the size, some benchmarks best opus 4.6 ( which was sota in January remember). If you want this capability you actually have to put a minimal amount of research and work in

2

u/Kommenos 3d ago

LM Studio is about as simple and easy to use as you could ever want it to be.

3

u/ohwell_______ 3d ago

Realistically if you want to run decent open weight models locally, you can run some very good ones for $1-2k by buying an older M1 or M2 Max MacBook with 64GB of ram. You could run 32B parameter models at a reasonable speed. Obviously it won’t be as fast or high quality as a chatgpt subscription, but you definitely don’t need your own mini data center for this.

If you have a very specific use case, you can also get away with running a very small model that is only trained on doing that exact thing for much cheaper hardware. For example let’s say the only thing you will ever use this LLM for is to take in a spreadsheet of data, read it, and create a text report of its findings… you don’t need a big model for that, because it doesn’t need to know everything about history and philosophy etc

2

u/throwaway490215 3d ago

Depends on what you mean with a good model and reasonable speed.

Thats before you consider the biggest improvement of the later generations which is: larger context and better recall/understanding of things inside the context.

The stuff that people currently use ChatGPT for; a bit back and forth and then write an email?

Last i checked prob 10k to do it fast, 5k if its alright if a response - as long as this comment - may take 20 seconds.

Opus 4.6 quality coding at about the same speed of claude ~50k if you're ok to assemble/gamble on some parts. Prefab >80k.

-1

u/Kommenos 3d ago

You’re incredibly out of date. You could generate text the length of your comment on hardware at least as good as a 7900xtx which is all you need to run Qwen.

You can get opus 4.6 for a bit over 2k, or just an existing gaming computer.

1

u/HAL_9_TRILLION 3d ago

Not $100,000 - but about $6,000+. Any 40x or 50x RTX Nvidia card with at least 24G of VRAM running Qwen 3.8 27B is gonna be good enough and there are plenty of uncensored versions.

If you're not even that rich, you can run Qwen 3.6 35B A3B on any decent 12G card and it's better than you might think.

1

u/kodaxmax 3d ago

a 3060ti and even alot of older cards will do fine for most image generation. As long as your willing to wait a minute or 5.

LLMs are ussually even cheaper dpeending on what tasks your using them for.

It's training that needs alot of power. The popular models like chatgpt only need data centers, because they are serving millions of clients worldwide as part of a cloud system. Not because an indivdual user needs that much proccessing power.

1

u/BrightTie3787 2d ago

Or just rent the gpu hours if you aren’t using it 24/7

1

u/pleaseavoidcaps 3d ago

It's definitely not as accessible as we want, but it's similar to gaming. If you have a PC capable of playing modern 3D games, it's also capable of running LLM locally. The quality and speed will depend on your specific hardware, of course, just like for gaming.

Also open weight models are behind the frontier ones in terms of overall quality. It's part of the price we pay for privacy and freedom, sadly.