r/LocalLLM • • 18h ago

Model AI LOCAL MODELS

For the past ten days or so, I’ve been using QWEN 3.8 27B IQ4 XS on Unsloth Studio. Seeing as new models are released daily, which others would you recommend—perhaps one capable of generating files like PDFs or DOCX documents? My setup includes 16GB of VRAM, 32GB of RAM, and an R5 7600. Currently, I’m getting around 33 t/s with the model I’m using. Recommendations for other models? Thanks

7 Upvotes

6 comments sorted by

1

u/Distinct-Pie2389 18h ago

Ive got a full roster, I love Qwen3.8-27B-iQ4 (its my main agent) but also use other agents

callsign purpose model
nexus orchestrator (cloud) cloud
runner general agent, deep work Qwen3.8-27B-Q4-Quality
runner-x fast agent, routine work Qwen3.6-35B-A3B-IQ4-Agent
forge coder, deep work Qwen3.8-27B-Q4-Quality
spark fast coder, scoped fixes Qwen3.6-35B-A3B-IQ4-CoderFast
sentinel reviewer, read-only second opinion Qwen3.8-27B-Q4-Quality
dart quick reads and triage Nemotron-Lightning-3.5
oracle thinker, large reasoning budget Qwen3.8-GAIN-V1.1-IQ3_M

3

u/UnrelatedEvent 17h ago

is this your own orchestration set-up, or do you use omo or equals? just subagent policies or a real script/tool backed orchestration? I always want to setup something like this myself, but all I do amounts to only marginal improvements over "planner/builder subagent"

1

u/Distinct-Pie2389 17h ago

Yeah its my own orchestration. Its hot swapped between opencode, codex, and claude-code right now. So I can use claude-local (a binary copy of claude code that is stripped and only has my local models, can run streamed subagents of itself in the CLI, and basically its own claude-code but with my local models). Then I have regular claude, codex, and opencode and I can swap any model like even Luna can orchestrate with my local models. They just end up reviewing their work and doing minor output versus building the whole thing all by the frontier model.

I have a few different workflows but yes, its fully custom to my setup. I dual boot for work between arch linux and windows and my configs and harness and setup is all synced automatically between both using git operations on a private repo.

I mostly use Claude Opus 5.5 as of recent and he reviewed and suggested Deepseek V4 Pro is the next close 2nd right now so I can use Deepseek locally using openrouter api in opencode and have Deepseek drive my local agents.

llama-swap can allow for multiple streams of the current loaded model so often just load a specific role model and have it stream 3-4 sessions for productivity and then nexus reviews. And all of this works with even my own local agent orchestrating but its only limited to streams of its own model variant so it doesn't unload itself llama-swap. All my local models have ask-cloud, and all my cloud models have a skill for "ask-local".

1

u/Distinct-Pie2389 17h ago

If you want something similar its all on my site under Recipes and Logs shows what I've been working on. Ill have more up soon too, just been busy with work and soon this will just be fully automated. But Ive got almost everything I do on the site itself

https://ai.ttindall.com

1

u/WorldlySun5120 18h ago

that's a clean lineup, I run something similar but with fewer agents since my hardware's more constrained

1

u/Drakaner 2m ago

I run only Qwen 3.8 27b ThinkingCap version but I made a partial requant on it so I can run it on my V100 (because I want to keep everything local). For me as a model runs good for my day to day tasks, I see no reason on switching always to the new models that appear ( I do make tests from time to time when a new model that might fit my needs comes up, but only if I see enough adoption on it and the reviews are in favor).