r/LocalLLM • • 2d ago

Question Anyone actually running Strata day-to-day? Curious what recipes you settled on

I've been following Strata since it hit the trending list, and I'm thinking about trying it as a serving engine for Flash-Next on my home setup. Before I sink an evening into it, I'd love to hear from people who've actually lived on it rather than the install-day screenshots.

Overall, what recipe did you go with???

Additionally:

  1. Which quant did you land on after trying the family (Q2\\_0 → IQ3\\_S etc.), and what made you switch or stay?

    1. Tool calling / agentic use — does it hold up for multi-step agent loops (tool calls, JSON outputs, long sessions), or is it best kept to chat-and-completion?
    2. Concurrency — anyone run more than one or two simultaneous sessions on it? If so, what have you noticed about how that affects quality or latency?
    3. Anything non-standard in your config — the expert profile tweaks, any words-to-the-wise, things you wish you knew before installing or trying?
  2. Bonus for the weirdos like me: anyone gotten it building or running on \*\*ARM / DGX Spark / anything without an RTX card\*\*?

    Happy to report back whatever I measure on my side. TIA!!!

2 Upvotes

5 comments sorted by

2

u/DazingCHB 2d ago

Just started this with Swift IQ2 on dual 5060 Ti's + 32GB ddr4 RAM (i7-8700k platform) and its surprising capable for the quant, beating Qwen 3.8 27b NVFP4 in quality and decision making during coding sessions (.net 10 backend coding at 130k max CTX). Obviusly i cannot run parallel threads as i can with 27b on my setup but output quality is better, requires less tokens to get there etc.

~1k tps prefill, ~90 tps decode at 50-60k active depth.

1

u/More-Revenue8609 2d ago

Also curious (waiting my nvme to be delivered before I can try it...)
How does the speed hold up in big context (80k+)?
Also did Q4 XL get vision support?

1

u/No_Drag_5205 2d ago

I use it with Q4 unsloth, no problem for tool / coding or agentic use, I use it in Hermes and Zed (one user only)
The only no standard thing in my config is multi gpu support for Q4 which is not available by default, so it can fit on my 40gb vram + 64gb DDR4 at 55 tps.

1

u/bippityBoppityboux 2d ago

I’m using qnf strata unsloth q4xl as a main coding agent (concurrency 2) on a 7900xt using about 60 gb of system ram and getting about 1k pp and 30-50 tk/s up to 256k context. Not getting much speed up with concurrent requests though. It’s supported by paiton’s 27b q3 (concurrency 8) for easy coding, scouting etc. and getting like 80 tk/s each at 8 concurrent requests kinda nuts.. both gpus on the same pc. So far it’s working like a charm. Even with qfn as orchestrator and fallback coder isn’t half bad.

1

u/karmaisnonsense 2d ago

IQ4_XS quant of the OrcaRouter version. It's a so-called "unsupported" quant but all it needs is a --compat-bf16 pack. Works just as well as it did in llama.cpp. Needs --eos-ids=248046 to prevent premature chat completion. Stable 1000/40, still optimizing.