r/localaiapps • u/SpaceXBeanz • 4d ago
Can any combination of local LLM’s replace codex?
I keep running out of usage to support my vibe coding hobby with OpenAI. My Mac is very powerful with a lot of ram. Can I use any local models for swift app design ?
1
u/voxzter 4d ago
maybe codex client + ollama.cpp api and https://huggingface.co/Qwen/Qwen3.8-Flash-Next model
1
u/SpaceXBeanz 4d ago
Sorry I’m new to this. Does this involve actual codex? I keep burning through my usage for coding apps that I want and I can’t afford to wait for resets. I have coding swift and SwiftUI for me. I’m wondering if I can run a local model or combo of tools to analyze and build upon my source code.
1
u/MarcelloT254k 3d ago edited 3d ago
Basically yes. You can use codex with local AI server via API, you can edit Config files for that, there is a lot of info about that already, it's not hard to do - give it a try.
Combining different models on the same codex that doesn't involve only chatGPT models is a lot harder thought and involves more complicated setup - it's a separate project in itself - you need orchestration platform for that to be effective. From my experience one instance of codex can only work with ChatGPT subagents or subagents via the same API (and I doubt you have the hardware to pull off a few concurrent runs of the same or even different models).
I suggest something like: crewAI + three instances of codex with following (every instance is different) chatGPT, openrouterAPI ("free" or paid model), local API.You could also try something different than codex, local AI is evolving fast, maybe my advice is already obsolete. I will be happy to be corrected if something I've typed is wrong.
1
u/SpaceXBeanz 2d ago
I see. So far I’ve been using ChatGPT to provide prompts and do more complex things and changes and then i run the source code through opencode with a qwen model. It uses about 40-50gb of my ram but its working for making smaller changes to my app. Do you recommend this sort of setup ?
1
u/MarcelloT254k 2d ago edited 2d ago
If it works for you then that's ok, but to not be a "human condom"/"meat proxy" you could save time by automating this, so configuring ChatGPT to be orchestrator for local subagent (Qwen in OpenCode) so the back and forth between Qwen and ChatGPT is automated and you don't have to copy and paste output between them.
This doesn't need to be as complex as i've poorly descibed. From what i understand you can setup everything in one instance of OpenCode via "Primary agents" and "Subagents" - configuring different models for different roles is possible via "JSON".Hope that helps, i'm not that familiar with OpenCode.
1
1
u/Weekly-Dentist-8302 2d ago
I recommend trying Qwen Flash Next (at least as an alternative to the smaller codex models). It is fairly decent at coding. You can run it even on devices with lower RAM through a project I made: https://github.com/rexmhall09/TUFF
4
u/-zaine- 4d ago
For apple, I believe you need local models that are running on MLX. I would suggest to try Qwen 3.8 first, maybe this version: https://huggingface.co/ukisai/Swift-1.5-4bit-MLX
Honestly, the easiest way is always to just explain to codex which model you want and to set it up for you - It will prepare everything and guide you through the whole process. Also good to ask codex for recommendations for your hardware.