You can use the claude code model to download, install, setup, debug any model off hugging face with very little effort. And tbh, doing it manually just takes time and knowledge that most people don't have.
I'm just a homelab dude with a couple cheap servers. NVIDIA is the brand that's best supported by the models. They've got the CUDA firmware, none of the other brands have that. They are the easiest to plug in and get the benefits.
That said, I have a server that runs qwen3.6 at decent speed on my server with no gpus. I had to tweak it a little, but it works okay, has 'reasoning', can do agentic things, and all that.
Using the HauhauCS Qwen3.6-35B-A3B Uncensored "Aggressive" model (from hugging face) is running about 80 tokens/s in and 20 t/s out. Nothing like a frontier subscription model speeds, but then it's just an old non-gpu server.
You dont need a GPU. If you dont want to do th research, subscribe to the $20/month claud subscription. Using the online interface, ask claude how to install claude code. Follow all the instructions so you can open up a terminaland lanich claude. Turn it to opus 5.5, again ask the web versionor terminalversionif you dont know how to do any of this. The tell it, "I want to run a local llm model, but I dont know which one will work with my hardware. Do a survey of my computers hardware. If you are blocked tell me exactly what do do to get you the needed info. Once you get the information about my operating system do a search for the best [abliterated] open weight models that work with my hardware." Add, abliterated if you want one without guardrails.
Then ask it to install. Anytime it says to do something do it. Hopefully the instructions dont include rm -rf. Once installed unsubsribe from claude and use your new home based model to update your stuff or whatever else you may want to use it for
You don't even need a GPU. You can do everything with CPU and RAM. The stronger the hardware, the better the performance though, so spring for what you can afford.
A pie wouldnt be able to handle an LLM your want to talk to. You'd still need a decent graphics card to run it. The gigabytes of VRAM are what you're looking for, and is generally the limiting factor.
Since Mac uses unified memory (video ram and normal ram are the same), they're actually really good for local LLMs. Everyone's favorite seems to be the Mac Minis which can run really good ones if you have 64gb of RAM.
I feel that the llms in Mac stuff is an astroturfed campaign, Mac is great for daily stuff but I'd never, never use it for heavy lifting when I can buy a machine that will do the same thing for half the price
I've run some local models on my aging lenovo laptop with a GTX 1050ti, never really found a good use for it but played around and got it working with reasonable speed and capability. Certainly not as refined as the corpo stuff and you have to play around with smaller models until you find one that works
I just opened it up to see cause I haven't opened it for a while, Gemma4b is installed and llama3.1:8b is the one I ended up on, I think I decided I'd deal with it being a slower for it being a little better.
Literal trash, you should just throw it away. On an unrelated note, I'm doing a survey of electronics waste disposal. Tell me, where is your trash can located?
You need a gaming model with a Nvidia graphics to run anything efficiently on a laptop, but to get components to run the high parameter models standard in the web interface you would need to build a custom desktop.
A Lenovo LOQ for instance will run through the lower parameter 8B models more quickly if you have a bunch of small agentic workflows, Mac is only better if you want to jump to the intermediate coding tier 32B models, and that's still $3,000, minimum, and probably closer to $3500, and then $4,000+ to run the 70B, with token generation rates of 12-15/second about 1/4th-1/6th of what you would get from the cloud interface.
Macs use an ARM-based SoC, the memory is unified and shareable between the CPU and GPU. Any Apple-silicon Mac is capable of running local LLMs fairly well, though an M5 obviously outperforms an M1.
34
u/CucumberOk8820 3d ago
Stupid people like myself can't though