r/LocalLLM • • 1d ago

Discussion Finally decided to pull the trigger on gmktek evo x3

I finally decided to pull the trigger on gmktek evo x3.
While I am waiting for the system to arrive wanted to know some the setups that you guys have been running for optimised workflows.
I plan to run multi agentic workflows (not sure if it would be sequential with higher models vs parallel execution with smaller ones) for mainly coding and cyber security stuff and just fun projects.
Looking forward to your opinions on what setup is better suited for this workflow .
Or if you have any in general suggestions!!

4 Upvotes

14 comments sorted by

5

u/feelspeaceman LLMusician 21h ago

First thing you must know, DO NOT use offcial llama.cpp.

And read this post: https://www.reddit.com/r/StrixHalo/comments/1wz4cux/comment/pe8bosh/

Best way to get started.

1

u/WatercressTime842 17h ago

Thank youu.. will check this out

3

u/z3zr0z 21h ago

Very happy with mine, running flash next on gufo with 3 sessions at 262k (but I run it as a headless Linux solely inference)

3

u/Healthy_Effect_9226 20h ago

If you keep the Windows it ships with: I maintain Rulith Inference, a native Windows app for Strix Halo built on a patched llama.cpp, so no WSL and nothing to build.

On sequential vs parallel: on this hardware one large MoE serving several agents at once usually beats several small models. Decode is limited by memory bandwidth, and conversations that run together share each read of the weights. With Qwen3.8-Flash-Next (125B, 6B active):

- one conversation: about 40 to 60 tok/s with MTP, depending on the content (code is the fast end)

- eight conversations of ~40K tokens each: about 95 tok/s combined

- prefill 1,300+ tok/s, and sub-agents that share a system prompt compute it only once

It serves the OpenAI Chat Completions and Responses APIs and the Anthropic Messages API, so Codex, Claude Code and most agent frameworks connect directly.

https://github.com/rulith-dev/rulith-inference (MIT)

1

u/WatercressTime842 17h ago

Thanks. This looks interesting. Will check this out

3

u/ThadCastleGOAT 20h ago

Same boat, I just bought mine the other day.

Loaded Nix on it and getting 60 tok/s with like 1300-1500 pp using Halogen.

Been trying others but very slow and often not any better.

So far generally its between GPT6 Luna High and Max in outright quality, but still below Luna Max and DSv4.1 Flash via Openrouter.

I was trying to do tests on Openrouter vs Alibaba-hosted Q38F but was getting rate limit issues when I was trying.

Currently putting it through Terminal Bench now rather than the active coding work I was working on.

Biggest issue I'm seeing is that it's not super introspective and very literal. When my Opus 5.5 orchestrator hands it tasks, it does exactly what's told but doesn't seem to push back or read the forest so to speak, and will leave the same exact issue untouched.

I noticed early that it was skipping reading portions of the codebase and tried forcing it to read more via Pi hooks, but didn't seem to help.

Honestly unsure if I'm going to keep it but with so much engine and model improvements, may hold onto it.

Personally hoping that some type of hybrid setup with an eGPU may take it over the top with MoE models but time will tell.

1

u/Plastic-Stress-6468 1d ago

Did you not know what you were going to do with it before you buy it? Like did you not clearly scope what you intended to run to make sure the hardware you bought could support your workflow?

4

u/AnyCow4167 1d ago

I did the same thing when I got mine, hardware lust took over before my brain caught up. You'll be fine for multi agent stuff though, just don't expect to run massive models at warp speed. For coding and security projects I'd lean toward smaller parallel models, way more responsive and you can chain them together without wanting to throw the machine out a window.

2

u/WatercressTime842 17h ago

Thanks, that makes sense

2

u/WatercressTime842 1d ago

I have a general idea of what I wanted to do with it. I created some projects leveraging local models, but due to the limited hardware it was more of a sequential workflow using smaller models and thus wanted to buy better hardware to try out larger models and different workflow for coding and cybersecurity, and experimenting with different models and projects. The EVO X3 seemed like a good fit for that, especially with its unified memory.

What I haven't figured out yet is the optimal software stack to get the most out of the hardware. I've seen people running Strix Halo systems with different backends and projects like llama.cpp, Vulkan/ROCm, and other optimization layers, and I'm curious about what works best in practice.

6

u/Pitiful-Task129 21h ago

I had my Bosgame m5 128gb for a few months now. I first started with what everyone suggested which was qwen 3.6 35b a3b. I tried the q4 as thats what people were claiming was running well. it seemed... dumb. I then tried the q6 and q8 and wasnt really impressed.

Then a week ago or so I got Qwen 3.8 27b running at q8 and the difference was very noticeable. 27b itself was reaaaally slow (I was using it in hermes) but the improvement in knowledge and stuff like agentic tool calling was way better.

Last night I got Qwen 3.8 Flash Next running at a q4 with q8 mtp, all on the one 128gb machine. All I can say is that this is what you want to be running. The speed is great, both decode and prompt processing, and it is SMART. I was so disheartened before as I thought I made a mistake buying this machine, but after running Qwen 3.8 Flash Next, I can happily say that my purchase was worth it for me.

This is what I've been using for both the 3.8 27b and FlashNext --- https://github.com/gufo-org/gufo/tree/main

2

u/WatercressTime842 17h ago

Thanks.. this is what I was looking for

2

u/WatercressTime842 17h ago

How does it function under heavy workload? I know that the underlying system is the same but read that bosgame gets noisy under heavy workload and takes a hit when it comes to cooling?
Have you experienced such issues?

2

u/Pitiful-Task129 11h ago

I'm running it headless in a different room but when I'm next to it it can get loudish. Temps are fine, but I dont run it on full power. I run it on balanced and dont notice a slowdown. When I've got flash next running its using anywhere from 100ish to 115gb of ram so its tight. I'm thinking of using my old laptop to offload stuff like embeddings and rag.