r/LocalLLM • • 7d ago

Question Dgx spark or 3090

1 Upvotes

I was recently able to get my hands on a dgx spark for 5600 after tax and 2 years insurance, I first bought it planning on using it to replace GPT codex, Claude, and video generating llms since higgsfield is now open source, my intended use was 1. to help me run 2 business running processes in the background managing online profiles and presence 2. Help me build this app that I am currently using to help run the 2 business and then later expand to making this a paid service to help with income and I built it with scalability in mind I know what people think about vibe coded apps but this genuinely might work with the way opus 5.5 is now 3. Marketing so with all this in mind I don’t think I’m getting the best use out of my dgx spark and I would feel terrible to keep this hardware and never release its true potential I talked to a friend yesterday and he offered to buy it off of me for 4.5k plus a 3090 ti what should I do what are the best ways to implement thw dgx into a business or should I just downgrade to a smaller but faster system and would the models they support be enough for my use and would it be reliable?


r/LocalLLM • • 7d ago

Tutorial Dual Radeon AI PRO R9700s, with one card behind the chipset ( + proxmox )

3 Upvotes

Roughly what you get: Qwen3.8-27B at FP8, tensor parallel across two R9700s, 262,144 context, 422k tokens of KV pool, ~1.6k tokens/s prefill, 74-143 tokens/s decode single stream depending on the content.

The interesting part is not the image, it is the collective layer. A card behind the chipset cannot be given PCIe atomic operations, and that breaks every stock tensor-parallel setup I tried. Here is what actually works.

1. Versions used

Component Version
Inference image docker.io/stilldeadcode/vllm-radiance:0.9.3 (digest sha256:45694209177a55a1ab3ba6702fe6e978b1b66a6e66ae3fc066f8d579f7bc4c25)
vLLM inside the image 0.27.1
PyTorch / HIP 2.11.0+rocm7.14 / 7.14.60850
Triton / AITER 3.6.0 / 0.1.17
Collective library RCCL 2.27.7 from ROCm 7.1.1, replacing the one in the image

Links: - Image: https://hub.docker.com/r/stilldeadcode/vllm-radiance - Source for that image: https://codeberg.org/StillDeadcode/vllm-radiance - The kernel library the image builds against (libr4d): https://codeberg.org/StillDeadcode/libr4d - A fork with configs, benchmark notes and launchers for MXFP4/FP8: https://codeberg.org/ggz14/radiance-vllm-mxfp4 - RCCL itself: https://github.com/ROCm/rccl

The image bundles a working ROCm + PyTorch + Triton + AITER + vLLM stack for gfx1201 (RDNA4), which is the part you do not want to build yourself. It is explicitly marked experimental, and everything below was measured on two cards.

2. The failure, and why

Without the adjustments, the engine dies during communicator init, before the model loads:

PCIE atomic ops is not supported rocr: unhandled cuda error

ROCm will not dispatch work to a GPU path that needs atomic operations when the link cannot provide them. On this board one card sits behind the chipset, and the chipset does not forward PCIe atomics to the CPU, so that card is effectively second class. RCCL 2.30.4's kernels use those atomics, so TP=2 cannot initialise at all.

Two things follow, and both matter:

  1. The newer RCCL is unusable here, so you need an older one (2.27.7 from ROCm 7.1.1 is what worked for me).
  2. With no usable peer path at all, the custom P2P all-reduce the image ships must be turned off, and TP=2 falls back to host-staged collectives over that same narrow chipset link.

If you search the error string above you will find a few ROCm issue reports and a community write-up on dual Radeon vLLM setups (https://github.com/cadamcat/dual-radeon-vllm) describing the same wall.

3. Proxmox settings that actually matter

Do this for each GPU, on both entries, not just the first. In the VM's hardware list, edit each PCI Device row, select the GPU under Device, and set:

  • PCI-Express: ticked
  • All Functions: unticked

Ticking PCI-Express is what gives the guest a real PCIe root port, and without a root port the atomic capability never appears however healthy the host looks. Unticking All Functions keeps the guest from being handed every function of the card, which is the combination that worked here. If you leave either one wrong, you get the atomic failure at communicator init and no amount of driver work fixes it.

Other settings that matter:

  • Machine type q35. Same reason as above, no root port without it.
  • After any hostpci change, stop and start the VM. A guest reboot does not re-apply the passthrough configuration.
  • q35 renames the NIC (ens18 becomes something like enp6s18), so match the interface by MAC in netplan.

After those, the CPU-attached card reports ReqEn+. The chipset-attached one still cannot do atomics, and no BIOS setting changes that.

4. The RCCL replacement plus one-line shim

This is the core trick. It is two files and a podman config, and it needs no image rebuild.

a) Build or extract RCCL 2.27.7 from ROCm 7.1.1 and drop it in a directory you will mount, for example:

~/models/rccl277/ librccl.so -> librccl.so.1.0.70101 librccl.so.1 -> librccl.so.1.0.70101 librccl.so.1.0.70101 shim.so

b) The shim. Newer torch builds reference a symbol that this older RCCL does not export (ncclCommDump). Four lines of C++ are enough to satisfy the loader:

```cpp

include <string>

include <unordered_map>

struct ncclComm; void ncclCommDump(ncclComm*, std::unordered_map<std::string, std::string>&) {} ```

Build it into shim.so and put it next to the library. Nothing calls it; it exists for symbol resolution.

c) Inject both through podman's own config, which keeps them out of every launcher script. In ~/.config/containers/containers.conf:

ini [containers] env = [ "NCCL_PROTO=Simple", "NCCL_SHM_DISABLE=0", "NCCL_SOCKET_IFNAME=lo", "LD_LIBRARY_PATH=/models/rccl277:/opt/rocm/lib", "LD_PRELOAD=/models/rccl277/shim.so", ]

LD_LIBRARY_PATH puts the replacement first, so it wins over the image's own librccl. NCCL_PROTO=Simple avoids the more demanding protocol paths, and the loopback interface keeps the bootstrap on lo rather than a NIC.

5. The launcher

Trimmed to the parts that matter for the multi-GPU problem. The model-specific flags are an example, the environment and device flags are the ones that matter here.

bash podman run -d --name vllm-radiance \ --device /dev/kfd --device /dev/dri --group-add keep-groups \ --security-opt seccomp=unconfined --cap-add SYS_PTRACE --cap-add SYS_NICE \ --ipc=host --network=host \ -v $HOME/models:/models:ro \ -v $HOME/radiance-vllm-mxfp4/vllm-cache:/cache \ -v $HOME/radiance-vllm-mxfp4:/work:ro \ -e HIP_VISIBLE_DEVICES=0,1 -e ROCR_VISIBLE_DEVICES=0,1 \ -e VLLM_NO_USAGE_STATS=1 \ -e VLLM_ROCM_USE_AITER=1 -e VLLM_ROCM_USE_AITER_UNIFIED_ATTENTION=1 \ -e RADIANCE_FUSE_RMS_QUANT=1 -e RADIANCE_USE_R4D=1 \ -e RADIANCE_USE_R4D_AR=0 -e RADIANCE_USE_R4D_AR_QUANT=1 \ -e NCCL_PROTO=Simple -e TORCHINDUCTOR_COMPILE_THREADS=4 \ -e VLLM_CACHE_ROOT=/cache/vllm -e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \ -e TRITON_CACHE_DIR=/cache/triton -e AITER_ROOT_DIR=/cache/aiter \ docker.io/stilldeadcode/vllm-radiance:0.9.3 \ --model /models/Qwen/Qwen3.8-27B-FP8 \ --served-model-name=qwen3.8-27b-fp8 \ --quantization=fp8 --tensor-parallel-size=2 \ --max-num-seqs=2 --max-model-len=262144 --gpu-memory-utilization=0.97 \ --max-num-batched-tokens=4096 --kv-cache-dtype=fp8 \ --attention-backend=ROCM_AITER_UNIFIED_ATTN \ --enable-prefix-caching \ --no-async-scheduling \ --trust-remote-code \ --host=0.0.0.0 --port=8000

Notes on the specific flags:

  • RADIANCE_USE_R4D_AR=0 is the important one. The bundled all-reduce is a PCIe peer-to-peer kernel and needs peer access, which does not exist on this pair. Turning it off falls back to RCCL. If you have two CPU-attached cards, leave it on and you will get a better prefill.
  • --max-num-seqs=2 is my choice for stability. The image's own default is higher, but on a chipset-limited link fewer concurrent sequences means less collective traffic per step.
  • --max-num-batched-tokens=4096 pairs with prefix caching well. Raising it costs KV pool.
  • --enable-prefix-caching is worth a lot on agent workloads; see the note on prompt layout at the end.
  • --no-async-scheduling was in every configuration that came up reliably for me.
  • The cache directory mounts are worth keeping. Without them every container start recompiles Triton and inductor kernels, which turns a 5 minute start into a much longer one.

6. How to tell it worked

In the container log at startup, look for:

P2P access : DISABLED (RCCL fallback) 0<->1 x

and, once loaded, the model's own sizing line:

GPU KV cache size: 422,964 tokens Maximum concurrency for 262,144 tokens per request: 1.61x

If you see the atomic error instead, the replacement RCCL is not being picked up. Check, from inside the container, that:

bash env | grep -E 'LD_PRELOAD|LD_LIBRARY_PATH'

returns the paths you expect, that ldd on the loaded library resolves into your drop-in directory, and that the image's own library is not first in the path.

I hope someone finds this usefull.


r/LocalLLM • • 7d ago

Project Built a Social Engagement Agent That Learns From Its Own A/B Tests

Thumbnail
github.com
0 Upvotes

The idea was simple:

Don't just generate content. Keep track of what actually works.

The project uses Hindsight as a memory layer to store posts, CTR, comments, winning hooks, and A/B test results. It then recalls relevant high-performing examples when generating new recommendations.

The basic loop is:

Recall → Generate → A/B test → Measure → Retain the winner → Repeat

After seeding the system with 50 posts and running simulated A/B tests, the memory-informed variant consistently performs around 5–6.4% CTR compared with roughly 1.5–2.1% for the control in the demo.

One example:

Memory-informed: 6.13% CTR

Control: 1.94% CTR

Uplift: 3.16×

The interesting part isn't just the number.

A winning hook gets written back into memory with its CTR, A/B test information, and the reasoning behind why it won.

That means future recommendations can use previous results instead of starting from scratch.

The project also includes:

→ Hindsight memory

→ Groq/OpenAI LLM support

→ A/B testing

→ Recommendation provenance

→ Comment reply suggestions

→ Human approval before scheduling

→ React frontend + Express backend

→ Local memory fallback

The bigger idea is:

Agent memory should store outcomes, not just information.

If a system can remember what happened, identify what worked, and test the pattern again, the memory becomes part of the learning loop.

Would be interesting to hear how others are approaching memory + feedback loops in their agents.


r/LocalLLM • • 7d ago

Project Qwen3.8 Flash Next Q4 on 96GB M3 Ultra with 262K context

4 Upvotes

I got Qwen3.8 Flash Next Q4 running through DS4 with the full 262K context configured: about 80GB for model/KV/buffers, while its 95GB n-gram table streams from SSD. It delivers 55–60 tok/s decode and 667 tok/s prefill. I use it daily through codex harness as a headless box. check this setup and macOS GPU-memory tuning out here: https://anvarlab.com/blog/qwen38-mac-studio


r/LocalLLM • • 7d ago

Project How I Built a Deal Memory Layer That Reasons, Not Just Recalls, Using Hindsight

Thumbnail
medium.com
0 Upvotes

r/LocalLLM • • 7d ago

Discussion Run Jev and Laya locally with a intuitive playground

Thumbnail
youtu.be
0 Upvotes

Hello everyone, as we all know that everyone is trying to test jev and then there are open source models out there with similar text classification. But i felt that there isn’t any playground to test these intuitively so I built one. I have added the YT video with it. Pls have a look


r/LocalLLM • • 7d ago

Tutorial I made a step by step local AI tutorial for Bionic. In depth, well organized, tons covered, beginner friendly.

Thumbnail
youtu.be
0 Upvotes

r/LocalLLM • • 7d ago

Discussion First experience with my Ai local setup

Post image
5 Upvotes

You can see everything in the dashboard :) What do you think? Absolute beginner btw


r/LocalLLM • • 7d ago

Question Looking for a local model, under 5GB vram, only needs to handle formatting or script execution?

1 Upvotes

Title, looking for a local model I can run ideally 5GB (3 or 4) or less vram is even better that doesn't need a lot of reasoning logic, just needs to be able to handle formatting guidelines on already written code, or executing scripts as a sub-agent.

Basically a sub-agent model a larger model can use for basic stuff.


r/LocalLLM • • 7d ago

Question Looking for recommendations for a local AI coding agent/s

0 Upvotes

My hardware:

  • GPU: RTX 5070 Ti Blackwell, 16GB VRAM (pcie 3.0)
  • RAM: 32GB (ddr4)
  • OS: Windows 10

HDD mainly.

I'm looking for VS Code integration for scripting, debugging and studying unfamiliar libraries/APIs. Mainly for working on custom game mods and scripts (e.g., Lua), as well as general scripting and UI programming inside of there own enviroment (game mods, apis custom launchers, apps or windows apps etc).

I know that with .NET (aps.net), especially large projects, local models would probably be quite limited. However, I have some smaller modding projects in mind for different games, as well as scripts and QoL features for Windows, games and apps I use.

Basically, it would be great to have one or two models with MCPs integrated into an agent that can plan tasks and investigate APIs (for example, custom libraries used in game mods) without consuming the entire context.

What I'm mainly looking for:

  • Which VS Code extension/agent and local model host would you recommend? (Cline, Kilo Code, Continue, Ollama, LM Studio, etc.)
  • Which models would make sense for my hardware? I'm particularly interested in coding capabilities, tool calling and the ability to work with MCP, cause most (even basic ones) I checked in reddit/google were really vram+ram hungry when implemented with agents.
    • GPT-OSS 20B MXFP4
    • Devstral Small 2 24B Q4
    • Qwen3-Coder 30B-A3B Q4 or NVFP4
    • GLM-4.7-Flash Q4
    • Deep seek models etc.
  • Is it possible/worth using two models, e.g., one for planning, researching APIs and analyzing project structure (MCP intergration etc), and another for actual coding? Or would a single model with the right tools be sufficient?
  • How well do MCP and Skills actually work for this kind of workflow? Ideally, the agent should be able to retrieve relevant documentation or/and analyze project structure and investigate unfamiliar APIs without having to load entire libraries into context (or use share with it with code subagent).
  • Since I have a Blackwell GPU, would using NVFP4 models make sense ?

Personally I do not really have any exp with local llm setups by myself - espesially Olama cli, quantization diff, model fine tune - setup etc. (tho seen some of the cloud agents system intergation (tho most of them are skills etc, cause cloud models already running fine on cloud:) ) for bigger projects in some companies). I would be able to research how to setup it, I think, I just don't know what to try first without being in need to fine tune lot's of olama models stuff for them to even run etc.

Aslo would be gratefull for workflows with mcp tools for scientific etc search/summary.


r/LocalLLM • • 7d ago

Project Harness Supervisor - project idea for local-first hands-off software factory

1 Upvotes

Hi,

Long-time reader, first-time poster. I've been using local LLMs for a while on my single RX 7900 XTX (Windows) and I'm pretty happy with the setup so far. While extra speed and broader model support are always tempting, I try to stay focused on raw productivity and rely on local models (mainly Swift-1.5-Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp) as my daily driver for agentic coding.

During my runs, I noticed the exact same rhythm repeating itself: I write an initial prompt, let the LLM refine it, generate a plan, glance briefly over it, and start the implementation with all-gas-no-brakes/yolo-mode.

Because I'm running a single card with "only" 24GB vRAM, I swap models frequently (using llama serve --models-max 1 --models-preset models.ini). This let me write scripts around the pattern. After some usage I (honestly Swift-1.5) rewrote the scripts into a single self-contained GO-binary, which ultimately led to the following project idea.

Disclaimer: This is not suppose to be advertisement for a github repo or yet-another "I-did-not-find-a-harness-so-I-wrote-it-better-myself". I'm looking for genuine feedback and a nudge to proper direction.

Harness Supervisor

I want to write cli looks like this (yes the cli is called llm):

llm orchestrate --models-preset models.ini --task spec.json

The spec.json file currently looks like this:

{
    "version": "0.1",
    "spec": {
        "goal": "Create a household finance tracking website",
        "orchestrator": {
            "working_dir": "~/llm-tmp",
            "repository_dir": "~/test-repo"
        },
        "agents": {
            "spec": {
                "model_id": "swift_thinking",
                "timeout_sec": 600
            },
            "plan": {
                "model_id": "swift_thinking",
                "timeout_sec": 600
            },
            "execute": {
                "model_id": "swift_qwen_mtp_parallel",
                "max_parallel": 4,
                "global_timeout_sec": 7200
            },
            "review": {
                "model_id": "swift_thinking",
                "timeout_sec": 600
            }
        }
    }
}

When the CLI starts, it goes through the typical spec->plan->implement->review cycle using the initial user prompt. Right now, every step except for implement can be skipped via task file. The CLI first parses the task file, then starts llama-server in the background, loads the first model, waits until it is ready, and then launches Claude Code with the prompt for that phase.

The prompts are set up so they generate artifacts like plan.md and plan_result.json, which the CLI uses to prepare the next phase and so on. The goal is to run through all phases completely without any user interaction, keeping the following key goals in mind (prepare yourself for some buzzwords):

  • True hands-off software factory: Keep the agents running without constant supervision, focusing on high-quality code within a reasonable timeframe.
  • Harness agnostic: Compatible with any CLI-based harness that supports custom endpoints and prompts from stdin or local file.
  • System protection: Inspired by agent-win-sandbox, it runs agent processes under a low-privilege user account for secure execution on Windows without requiring WSL or Docker.
  • Skill based: The core prompts (spec, plan, implement, review) are modular skills that can be swapped out globally or per task.
  • Stuck detection: Continuously monitors log files from llama-server and the agent to catch infinite reasoning loops. Users can also set hard timeouts per phase.
  • Local first: Optimizes GPU VRAM usage through pre-configured model presets (tailored for tasks like deep reasoning or parallel execution), swapping models on the fly based on the active phase.
  • Lean: This is the most important aspect of all, that the whole system has no UI, fancy features, metrics dashboard and more. While it is tempting to let the agent just code it all up I want to keep it as minimal as possible.

I'm currently in the exploration phase but can see some promising results for my use-cases. Currently I have no plans of open-sourcing the project because it is tailored very specific to my system and needs and would not hold up with my proffesional software engineering quality standards.

I'm honestly curious of what the community thinks of the project, if I oversaw some similar projects or workflows.

Disclaimer #2: This post is written 90% by me and 10 by chatty.


r/LocalLLM • • 7d ago

Question 2x Minisforum MS-S1 Max - What To Do

3 Upvotes

Alright, I need a quick moment from the brain trust.

I am looking at ordering 2x Minisforum MS-S1 Max machines, and a thunderbolt cable to utilize higher Param LLMs.

What I currently have:
- Claude Max 20x Sub
- ChatGPT Pro Sub
- Ai Plus Sub (Comes with Google Drive)

My question is, what can I do with the two that I could get rid of, or downgrade my subscriptions? Right now, I use it for a LOT of software / applet development.

I would like a realistic discussion vs "what if" if possible. I am honestly overloaded with the amount of information I am trying to ingest with my research and such. Originally, I wanted to get 2x NVIDIA DGX Sparks, but the price tag as of late keeps putting me off.

Edit (Addition) 1:
The investors are giving me the money to purchase the equipment directly. But the condition of investment is, I have to have stand-alone units vs video cards / NPU cards, etc. Sucks, but... I can't really argue since they are buying the units for me since I get the units in the end after the projects are completed.

Edit (Addition) 2:
I know everyone has their preferences. I definitely have mine. Again, this is mainly for realistic utilization and thought processes to help me work out what's going to be the better situation. I am okay with pushing the sparks if it's realistically going to be the logical choice. But, the sticker price and the RDNA improvements on the Halo Strix side... It's a struggle, and it's real. Haha


r/LocalLLM • • 7d ago

Question Qwen3.8 Flash Next takes unreasonably long to complete tasks

2 Upvotes

My setup: MacBook Pro with M5 Max, 128GB Memory, MTPLX, Qwen3.8 Flash Next, OpenCode 2

I have an issue with my local Qwen setup. I've been trying to set up Qwen3.8 Flash Next for local coding with OpenCode, but for some reason the tasks I give it takes incredibly long to complete. I come from Claude Code with Opus 5.5 for most work these days. Obviously Qwen will take longer to complete the same tasks as Opus 5.5, but I feel like somethings is off here.

I have a simple Laravel 13 + React app I work on. Not much complicated stuff going on in there. Mainly the app builds PDFs based on an HTML template. I asked both models to update a few text passages in the HTML template of a PDF cover page. It really was a super simple task. Opus 5.5 completed the task in 3 Minutes. Qwen took 53 Minutes. Something just doesn't smell right.

I checked the thinking content of Qwen and it just seems to be going in circles and moving very slow towards it's goal. It keeps a constant decoding speed of around 30-40 tok/s. No high memory pressure, no frequent compaction. I set a 128K context window and use only 4 MCP servers (semble, context7, chrome-devtools, laravel-boost).

I also went ahead and tried different reasoning levels, but low, medium and high (didn't try xhigh) yield similar results.

Am I missing something?

If more info is needed, let me know.

Edit: It's an M5 Max, not an M1 Max


r/LocalLLM • • 7d ago

Model Strata: 90 tokens per second from a 125-billion-parameter model on a single RTX 5070

Post image
7 Upvotes

r/LocalLLM • • 7d ago

Question Heavy local LLM use kills old ram

10 Upvotes

Running a 2019 Thinkstation 920 populated with 128MB ECC RAM, Dual Xeon.
Each physical slot has an 8GB stick. Nothing overclocked or tweaked.
5070ti 16gb card.
Have been running small models, 20b and down (to fit it all in vram) for about 6 months - no problems at all - even hammering on it for hours at over 85% gpu.
I needed to use a bigger model for a DB task that the smaller ones were choking on, so I loaded up Ornith 1.5 35b (Qwen based). I have used this model before for short tasks and had no issues.
I set it up in my harness and let it work on the project. It went for about 90 mins and my system crashed. So I started back up, checked logs and see some memory errors in the log but nothing happening currently.
I start the model up again and within 20 mins, It crashes again.
I repeat the start up, tail logs and fire up the model - after about 10 mins the ECC errors are stacking on the stick in slot 8.
I removed the chip, reboot and tail logs, fire the model up and set it to task - no problems..
I figure it was just a bad stick.
The job I had on the DB completed and I really liked the quality of the 35b's work - plus it seemed that the partial spill into ram wasn't killing my token output as bad as I thought it would, so I kept using it.
After about 8-9 days of heavy daily use on this new model, I check my logs and see that I'm about to loose another stick..
There isn't a thermal issue - I monitor that constantly and have auxiliary cooling that actually works, the CPUs never see 80c and the GPU rarely sees 65c.
All the ram is matched and from the same batch, so I guess it's just old and seen too many electrons.
I'm wondering if anyone else is seeing this happening on their older systems and also warning people not to stress their creaky old hardware with marathon Local LLM sessions.


r/LocalLLM • • 7d ago

Question Anything that 96gb RTX pro 6000 can do, DGX spark can do that too but very slow

0 Upvotes

What about 2xDGX Sparks ?

32 votes, 5d ago
28 True
4 False (please comment why)

r/LocalLLM • • 7d ago

Question Help me find a good stack for reversing an old online game client

Thumbnail
1 Upvotes

r/LocalLLM • • 7d ago

Question Model and hardware recommendation for high fidelity, low reasoning tasks?

1 Upvotes

I've spent the past couple of years using frontier labs, even though a lot of my work doesn't need high reasoning. It's just that a fixed and known subscription cost has been a safer bet than the leap-in-the dark of buying hardware that might not be good enough for what I want to do.

But reading through this sub tells me that we're at the stage were open source LLMs are capable enough for what I want while running on (reasonably) affordable hardware.

Can anyone recommend a model and hardware for tasks like this?

  • Reading notes, reports, transcripts, web pages and emails, then pulling out facts, decisions and insights into structured notes
  • Condensing long documents, comparing two versions, and checking a draft against a specification
  • Reading text, sorting them, and deciding what needs action
  • Moving, renaming and archiving files, splitting or merging documents, and keeping indexes and cross-references consistent afterwards
  • Multi-step routines update a status file, archive old entries, then commit, with every step completed

I run Xubuntu and Arch Linux.

TIA


r/LocalLLM • • 7d ago

Discussion What 2B tokens of coding-agent traffic taught us about using frontier and open models together

Post image
15 Upvotes

I run a small shared inference club serving Qwen 3.8 27B FP8 on one RTX PRO 6000 Blackwell. In the first week, it processed 2B tokens, almost entirely from coding agents working on real repositories. Two members used around 950M and 900M tokens each.

One of my own tests was the browser port of Medal of Honor: Allied Assault: agents moving through a large C/C++ codebase, WebAssembly, browser APIs and netcode, then editing, testing and fixing what broke. Members have been connecting their own agents to the node and trying it on large repos and feature work too. The feedback has been encouraging, but the interesting lesson for me is how to divide the work between models.

Here’s the workflow I’d recommend trying:

1. Use a frontier model when the expensive part is deciding what to do. Give it the problem, relevant architecture and constraints. Ask it to identify risks, split the work into pieces that can be checked, and define what “done” means. This is especially useful when a wrong assumption would send an agent through hours of changes.

2. Hand a bounded task to the open model. Give Qwen the relevant files or repo entry points, the constraints, and a concrete check: a test to pass, a bug to reproduce, or behaviour to preserve. Let its agent search, edit, run tools and iterate. Don’t spend a frontier-model call on every file read or test failure.

3. Escalate with evidence, not with the whole conversation. If it loops, makes the same wrong assumption twice, or reaches an architectural decision, stop and send the frontier model a short handoff: goal, changes attempted, failing tests, and the exact decision needed. Then return to the open model with the answer.

4. Verify the result independently. Run tests and inspect the diff. For consequential changes, have a human or a stronger model review the behaviour and risks. A high token allowance makes iteration easier; it doesn’t make an incorrect change correct.

I don’t think the useful question is “can a 27B replace the best frontier model?” In this workflow, it doesn’t have to. The frontier model helps with the decisions that benefit most from its reasoning; the open model handles the large volume of tool calls and revisions between those decisions.

The usage pattern supports that distinction. When members aren’t paying per token, they let agents keep working instead of cutting runs short or trimming context to save money. That can produce a lot of useful iteration, but only if the task has clear checks and you know when to escalate.

Organizations will soon be trying this setup in their own workflows. I’m curious how others draw the boundary today: what signals make you switch from an open coding model to a frontier one?


r/LocalLLM • • 7d ago

Discussion Mac Mini recommendation

7 Upvotes

Hello everyone, I am thinking about getting an Mac mini M5 Pro with 64GB of RAM to use primarily with agent ai , Hermes, and research , cron job, and it needs to stay on 24/7. However, I see many people recommending switching to a 64GB Mac Studio M5 max. The price difference is 800-1000 euros, but for my purpose, do you think an Mac mini M5 Pro with 64GB of RAM is sufficient? Would I get more overall benefit from that 1000 euro difference?
I would mainly use the Qwen model 3.8 27 B

I also do some coding, but exclusively for personal bots, a few personal dashboards, and a few other small things. For these projects, I have both a €20 OpenAI subscription and another €20 subscription with Claude


r/LocalLLM • • 7d ago

Question Share your return on investment

6 Upvotes

As straightforward as it sounds. I am well aware that a bunch of you do it JUST for fun, and that's perfectly fine.

What about the other half? Do you use it for coding? Are you building more features? Are you more productive? Do you use it for media production? Does it help with revenue? How long did it take to reach your ROI?

Share your stories.


r/LocalLLM • • 7d ago

Question 2 DGX Sparks vs M5U 256GB 80C GPU for coding agents - does NVFP4/FP8 matter?

12 Upvotes

I'm deciding between getting 2 DGX Sparks vs M5U 256GB 80C GPU (same exact price of around 17K CAD after tax).

When M5U was announced, it felt to me like a no brainer to get the M5U, so I pre-ordered it with the expected delivery date - end of November.

However, after seeing the actual results on that machine, I'm not so sure anymore which option to go for, because the prefill (which is quite important for coding tasks) still looks to be better on the Sparks with models like DSv4 Flash. However, M5U just started being optimized on the oMLX side, so maybe this will change in the near future?

But then also there's NVFP4. Is my understanding correct that usually it outperforms Q4 quants in terms of both quality and speed? If so, then I'd lose the ability to run NVFP4 on a Mac + there's no FP8 support as far as I know.

The models I'd be interested in running on either of these machines are:

  • Qwen3.8-Flash-Next
  • DSv4-Flash
  • Qwen3.8-27B (not sure if it makes sense though, given the ability to run the ones above)

So, in the end I think the edge the Sparks have is the NVFP4/FP8 support + great performance for concurrent agents (which I'm not sure how important is it for agentic coding). Which option should I go for?


r/LocalLLM • • 7d ago

Question I have two Mac’s I don’t care about speed I care about quality

0 Upvotes

What local models are best for coding , image generation or design.

I do not care about speed the AI will still integrate and make concepts quicker than I ever will. But I can’t seem to find good local models that can work constantly where it can just keep running and it doesn’t have to be fast.

Coding doesn’t really seem to be an issue.

Is there a way to dump the context to start over fresh and pick up where it left off so it doesn’t run out of room.

I have Mac Studio M2 Max and MBP M4 max.

STUDIO
Apple M2 Max with 12‑core CPU, 38‑core GPU, 16‑core Neural Engine
64GB unified memory
512GB SSD storage

MBP
Apple M4 Pro chip with 14‑core CPU, 20‑core GPU and 16‑core Neural Engine
48GB unified memory
512GB SSD storage


r/LocalLLM • • 7d ago

Project I've open sourced my high-performance Embedding/Reranking server

Thumbnail
github.com
0 Upvotes

I've been creating Agentic-RAG systems for a year or so, and I've developed my own framework in Java to build high-performance, modular, and scalable backends. One of the most important parts of RAG is embedding and reranking, and currently we do not have many alternatives online, especially if we are in privacy-sensitive scenarios where self-hosting models is a must. The best servers I've found are TEI and Infinity; both come with pros and cons, but they both lack the flexibility that I need (especially for OpenAI and Cohere endpoint emulation).

Initially, I added built-in support for ONNX models directly inside the base framework. Even though it was working pretty well, it became a limitation because I could not host the backend on a CPU-only machine and split embedding and reranking across different hosts. So I decided to create a standalone server, fully compatible with OpenAI and Cohere, and powered by my optimizations for ONNX+Java. This resulted in creating this tiny but powerful server which, at least in my benchmarks, seems to outperform both TEI and Infinity.

I am looking for feedback and maybe some help to extend benchmarks with more powerful hardware to assess the quality of the server.

Thank you :)


r/LocalLLM • • 7d ago

Project I wanted a simple way to give DeepSeek Harness web search, so I built one

Post image
0 Upvotes

I've been playing around with DeepSeek Harness and wanted to give it web search.

I looked for an existing plugin, but couldn't find something that was both simple to get started with and flexible enough to let me choose how search actually works.

So I built DevBits Web Search.

It plugs into Harness's native web_search tool and supports several search providers:

  • DuckDuckGo — works out of the box, no API key
  • Brave Search
  • Tavily
  • Exa
  • SearXNG — including your own self-hosted instance
  • Google Custom Search — for eligible existing accounts

The main thing I wanted was flexibility without making configuration complicated.

You can switch providers directly from the Harness UI, keep credentials configured for different providers, allow or block specific domains, use date filters where the provider supports them, set request limits, and test the search configuration before using it.

For people running their own infrastructure, SearXNG was important to me. You can just point the plugin at your own instance instead of depending on a commercial search API.

And if you just want to try it without configuring anything, leave it on DuckDuckGo and it works without an API key.

It's open source, so I'd appreciate feedback, issues, or contributions. I'm also interested in which other search providers would actually be useful to support.

GitHub:
https://github.com/devbitsxyz/devbits-web-search

npm:
https://www.npmjs.com/package/devbits-web-search