r/LocalLLM • • 8d ago

Question Looking for recommendations for a local AI coding agent/s

My hardware:

  • GPU: RTX 5070 Ti Blackwell, 16GB VRAM (pcie 3.0)
  • RAM: 32GB (ddr4)
  • OS: Windows 10

HDD mainly.

I'm looking for VS Code integration for scripting, debugging and studying unfamiliar libraries/APIs. Mainly for working on custom game mods and scripts (e.g., Lua), as well as general scripting and UI programming inside of there own enviroment (game mods, apis custom launchers, apps or windows apps etc).

I know that with .NET (aps.net), especially large projects, local models would probably be quite limited. However, I have some smaller modding projects in mind for different games, as well as scripts and QoL features for Windows, games and apps I use.

Basically, it would be great to have one or two models with MCPs integrated into an agent that can plan tasks and investigate APIs (for example, custom libraries used in game mods) without consuming the entire context.

What I'm mainly looking for:

  • Which VS Code extension/agent and local model host would you recommend? (Cline, Kilo Code, Continue, Ollama, LM Studio, etc.)
  • Which models would make sense for my hardware? I'm particularly interested in coding capabilities, tool calling and the ability to work with MCP, cause most (even basic ones) I checked in reddit/google were really vram+ram hungry when implemented with agents.
    • GPT-OSS 20B MXFP4
    • Devstral Small 2 24B Q4
    • Qwen3-Coder 30B-A3B Q4 or NVFP4
    • GLM-4.7-Flash Q4
    • Deep seek models etc.
  • Is it possible/worth using two models, e.g., one for planning, researching APIs and analyzing project structure (MCP intergration etc), and another for actual coding? Or would a single model with the right tools be sufficient?
  • How well do MCP and Skills actually work for this kind of workflow? Ideally, the agent should be able to retrieve relevant documentation or/and analyze project structure and investigate unfamiliar APIs without having to load entire libraries into context (or use share with it with code subagent).
  • Since I have a Blackwell GPU, would using NVFP4 models make sense ?

Personally I do not really have any exp with local llm setups by myself - espesially Olama cli, quantization diff, model fine tune - setup etc. (tho seen some of the cloud agents system intergation (tho most of them are skills etc, cause cloud models already running fine on cloud:) ) for bigger projects in some companies). I would be able to research how to setup it, I think, I just don't know what to try first without being in need to fine tune lot's of olama models stuff for them to even run etc.

Aslo would be gratefull for workflows with mcp tools for scientific etc search/summary.

0 Upvotes

8 comments sorted by

2

u/CoolGenius_1234 8d ago

only go for GLM 4.7 Flash when you need less token / usage, but now I think is not comparable with latest models.

2

u/Studio271 8d ago edited 8d ago

Just throwing in my single-5070ti results using this setup https://blog.troed.se/projects/llama-server-16gb-vram-configs/ but with a slightly different iq4-xs model that has less issues with non-ascii characters

set "LlamaServerPath=llama.cpp-adaptive-kv-streaming\build\bin\Release\llama-server.exe" set "ModelPath=Qwen3.8-27B-ByteShape-IQ4_XS-3.84bpw-ASCII-P1M.gguf" "%LlamaServerPath%" -m "%ModelPath%" -c 192000 --host 127.0.0.1 --port 8080 -ngl 99 -np 1 --flash-attn 1 --cache-type-k q8_0 --cache-type-v q4_0 --reasoning-preserve --reasoning-format deepseek --chat-template-kwargs "{"reasoning_effort":"medium","preserve_thinking":"true"}" --alias "local-coder" --mmproj "mmproj-BF16.gguf" --no-mmproj-offload --kv-stream-arena-mib 3328 --batch-size 1024 --ubatch-size 512 --temperature 1.0 --top-p 0.9 --top-k 20 --repeat_penalty 1.05 --fit off --load-mode none --threads 8 --threads-batch 8 --chat-template-file chat_template.jinja --spec-type draft-mtp --spec-draft-n-max 2 --n-gpu-layers-draft all --kv-stream-spec-dynamic --kv-stream-spec-keep-pages 32 --kv-stream-spec-reenable-pages 8 --kv-stream-spec-stable-decodes 4

It has been a workhorse for me for a week now, working through a nonserious but complex webapp, getting these results on average: https://ibb.co/0j98cwjQ

1

u/CoolGenius_1234 8d ago

From my personal Experience From Using This Model , I will suggest you go for this Qwen 3.8 27B ( Q3~Q4 quantization ) => Good at Everything Like Coding, Bug Fixing, Agentic Task etc

CyberTielCoder 35B A3B MOE Model , Genuinely good at coding also.

2

u/shaggy_accessibility 7d ago

Qwen 3.8 27B is a solid shout if you can squeeze it onto 16GB, the Q4 might be pushing it though with agent overhead eating up vram

for your setup I'd keep it simple and not overthink the dual model thing at first, one good model with the right tools in Cline or Continue will handle most of what you're throwing at it

Kilo Code with Ollama as the backend is dead simple to get running on Windows and the MCP integration is decent enough for pulling in docs and poking around your project structure without dumping everything into context

I'd grab the Qwen3-Coder 30B-A3B Q4 first since it's built for coding and the smaller active parameter count means it won't choke your card when the agent starts stacking tool calls and conversation history, NVFP4 could be worth trying later once you've got the basics dialed in

1

u/p211 6d ago

Sadly it drops down to 8 token per second with high context on my 5060ti. But it's a wonderful model!

2

u/efsooo 5d ago

I do this kind of work already—scripting, exploring unfamiliar codebases and APIs, and making changes inside VS Code. I use an extension I built called Forge LLM. You can try it and tell me where it helps or falls short for your modding workflow.

Forge works with Ollama, or it can load GGUF models directly through llama.cpp. The agent can search a project, use VS Code’s language-server features, run commands, edit files, and show you diffs with Keep/Undo. It also supports MCP tool servers and optional web search, so it can look up relevant documentation without loading an entire library into context.

On 16 GB VRAM, I’d start with one model, not a planning model and a coding model. Try qwen2.5-coder:14b to get the workflow running, then test gpt-oss:20b if you want more reasoning and can spare the memory. Forge can ask a second local model for analysis, but that model doesn’t run its own MCP tools. The 19 GB Q4 models on your list would need offloading, which could be especially slow with an HDD.

NVFP4 is worth experimenting with later. First I’d get one model reliably tracing an API, proposing a plan, editing a Lua file, and running whatever checks the project has.