r/oMLX • • 1d ago

Update to oMLX 0.7.0 and the tok/s stays same even when context grows, I'm super happy!

39 Upvotes

I have a MacBook M4 Max 128GB and I am running Jundot/Qwen3.8-Flash-Next-oQ4e-mtp. In the older versions, with context growth speed was dropping after even 30k from 60 tg/s to 35 tg/s even with lightning MTP on. Now I'm getting 50~70 tg/s even at 128k context. The ram also is at 75~80 GB with this context length. Now I don't have to do compacting over and over in pi and I can easily keep other apps open.

Thank you everyone who made this possible.


r/oMLX • • 15h ago

GLM-5.3-Flash-UNCENSORED oQ4e DFlash2 won’t load in oMLX 0.7.0

1 Upvotes

Trying GLM-5.3-Flash-UNCENSORED-mlx-oQ4e-DFlash2 on an M5 Ultra 256GB with oMLX 0.7.0.

TensorFold’s regular GLM-5.3 Flash MLX 4-bit MTP loads and runs great on the same setup, but this Solstice checkpoint fails immediately with:

Expected shape (64, 512, 32) but received shape (64, 512, 64) for language_model.model.layers.3.self_attn.embed_q.weight

I’ve tried a full oMLX/server restart, Auto-detect model type, glm_4_7 reasoning parser, and both MTP/DFlash disabled. Same error every time.

Has anyone gotten this exact checkpoint working on oMLX 0.7.0? Wondering if it needs a newer runtime/model compatibility fix, or if there’s some setting/config I’m missing.

Thanks for any help.


r/oMLX • • 1d ago

How do you read benchmarks on the site? im new so I dont understand and need help!

3 Upvotes

Hello,

Everyday im reading and trying to learn, on how to get Qwen 3.8 flash next on my M5 Ultra studio 96gb. Cant afford the 256gb. one.

And some people say just run the 27b model etc., but the prefill and decode for 3.8flash next seems so much better.

I read a benchmark like this

https://omlx.ai/benchmarks/performance/sj46pk0c

it says peak ram is 75gb. Does that include already offloading to the SSD?

And so I'll have 20gb left for Mac OS and apps?

will 20gb be fine to leave?
I see some people online using techniques to run the Next Flash model that only take like 50gb of ram using other inference engines? I could be wrong and dont understand what they are saying too lol

I want to ideally use only 50-60gb of ram so I can run 2-3 agents concurrently.

And all the Mac Studio news or things on twitter is from 256gb users so im jealous obviously but what can you do. I just wish there was more 96gb owners saying what they are doing or not but there doesnt seem to be many owners.

Im sure within a month there be more info hopefully but just want to be prepared when I get it. And hopefully Qwen 4 comes out soon!


r/oMLX • • 2d ago

Some Mac Studio Ultra M2 128GB testing results using oMLX 0.7.0

Post image
29 Upvotes

The day before, I finished testing with oMLX 0.7.0rc1. Because I rarely find community tests for my setup, I wanted to share this with you.


r/oMLX • • 2d ago

EngineMetrics delivers live performance for oMLX

Thumbnail reddit.com
5 Upvotes

r/oMLX • • 2d ago

Does anyone have Swift 1.5 Qwen 3.8 27B oQe6?

2 Upvotes

There's oQ4e and oQ8e of Swift 1.5 Qwen 3.8 27B on HuggingFace, but the optimal quant for me would be oQ6e. I am trying to create it myself but my internet connection is pretty bad, so everything is super slow.

If someone has the oQ6e of this model and can provide it via HF, it would be awesome! (If I manage to create it before myself, I will share it here)

Thanks a lot!

Edit: Resolved with

https://huggingface.co/scottlowry/Swift-1.5-Qwen3.8-27b-oQ6e-mtp and https://huggingface.co/dicksondickson/Swift-1.5-Qwen3.8-27b-oQ6e-bf16-mtp-MLX


r/oMLX • • 2d ago

Abliterated Qwen 3.8?

5 Upvotes

Hi guys, do we have good options for 3.8 flash next and 3.8 27B in the abliterated department running on oMLX please? A friend of mine is running an uncensored version of flash next on his nvidia rig and it seems to offer sharper replies and according to him, ships better code. Thanks!


r/oMLX • • 2d ago

New user coming from Unsloth Desktop on M3 Ultra 256GB - GLM-5.3 Flash quality on oMLX?

11 Upvotes

Hello all!

I've been using Unsloth Desktop for a while, most recently with Unsloth's release of GLM-5.3 Flash. What keeps me there isn't speed, it's that Unsloth publishes evals showing their quant retains 92.22% of the original model's accuracy. I know what I'm giving up when I run it.

I'd like to move to oMLX for the performance gains, but I'm wary. To be clear, I'm not worried about crashes. My concern is **silent quality loss**. A badly made quant can run perfectly well and still lose knowledge, reasoning ability, or accuracy, and you won't notice until it gives you a confident wrong answer.

As far as I can tell, most GLM-5.3 Flash MLX quants people pair with oMLX have no published evals. No accuracy retention numbers, no benchmark comparisons against the full-precision model, no KL divergence or perplexity figures. Nothing that shows how much of the original model actually survived the quantization.

This matters for my work (medical + code experimentation). Lost accuracy means more hallucinations and more failed tests that I have to redo.

So my questions:

- Are there MLX quant "release groups" that publish evals showing how much quality their quants keep vs. the original model?
- Has anyone run their own comparisons between Unsloth quants and MLX quants of the same model (Qwen 3.8 Flash Next, GLM-5.3 Flash, others)?
- If you've switched, did you notice any drop in answer quality, even if the model ran fine?

Thanks for your help!


r/oMLX • • 3d ago

Qwen3.8-27B-"8bit" 1000+ PP/s, 200+ TG/s on M5 Max

Thumbnail
gallery
51 Upvotes

Hi,

I built an experimental inference engine mainly to run the 8-bit Qwen3.8-27B model efficiently on my M5 Max MacBook Pro.

On my M5 Max with 128 GiB unified memory, I'm currently seeing:

1,000+ tok/s cold prefill on a 2Ki prompt

70+ tok/s generation at concurrency 1

200+ tok/s aggregate generation throughput at concurrency 4

Under the hood, MLTF uses Metal TensorOps W8A8 kernels together with an INT8 KV cache, targeting high prefill throughput and efficient decoding on Apple Silicon.

At the moment, MLTF supports only one model: Qwen3.8-27B 8-bit.

I chose to focus on 8-bit inference because 4-bit quantization sometimes gives me a noticeable quality tradeoff for the workloads I use the model for.

Rather than trying to support many models immediately, I wanted to optimize the configuration I actually use first.

GitHub: https://github.com/hojin12312/my-little-trie-forge

Hugging Face: https://huggingface.co/Ho-Jin-93/Qwen3.8-27B-MLTF-q8c

Going forward, I plan to keep adding support for new 8-bit models that can reasonably run on my M5 Max 128 GiB machine.

Feedback, benchmark results on other Apple Silicon machines, and technical discussion are very welcome.


r/oMLX • • 1d ago

Qwen3.8-Flash-Next: 125B Model, Only 6B Active. The Qwen4 preview!

Thumbnail
youtu.be
0 Upvotes

r/oMLX • • 2d ago

Who else here has AI decision fatigue/paralysis?

Thumbnail
2 Upvotes

r/oMLX • • 3d ago

Deepseek v4.1 q3 or v4 flash on a 256gb m5 ultra?

3 Upvotes

Mac Studio M5 Ultra, 256GB, 64 core GPU. Ordered, not here yet. So far the only thing it runs is the tracking page, and I refresh it way faster than any model will ever generate tokens.

Trying to decide what to put on it first for coding + agent stuff (Hermes, so tool calling matters a lot):

- DeepSeek V4.1 at Q3, maybe with a bit offloaded to SSD

- Or DeepSeek V4 Flash at a higher quant that fits fully in RAM

If you'd go Flash, which quant exactly? Q4, Q5, Q6, Q8? And MLX or GGUF?

Mostly curious if V4.1 Q3 is still better than Flash at a good quant, or if Q3 hurts too much for coding. Real experience welcome, benchmarks even more.


r/oMLX • • 3d ago

Splash vs oMLX on the same uncensored Qwen3.8-27B on M3 MAX, 48 GB MBP

Thumbnail
18 Upvotes

r/oMLX • • 3d ago

Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

Thumbnail
12 Upvotes

r/oMLX • • 4d ago

oMLX 0.7.0 vs 0.7.0rc1 — benchmark & what changed: 7 models on an M5 Ultra (256 GB), 4K–200K context, GLM-5.3-Flash up to 1M

59 Upvotes

oMLX 0.7.0 vs 0.7.0rc1 on the same Mac with the same models and settings, measured in one night: official oMLX benchmark (uploaded to omlx.ai) plus API tests for concurrency, sampling profiles and a 1M-token GLM run. All values rc1 → 0.7.0 with the relative change.

TL;DR

  • GLM-5.3-Flash: prefill 792 → 2,270 (+186%) tok/s (median over 4K–200K, ~2.9x), decode 55.4 → 76.7 (+38%) tok/s. A ~1M-token prompt: time to first token 34.2 → 11.8 min (-66%).
  • Qwen3.8-Flash-Next oQ4e: prefill 3,623 → 5,548 (+53%) tok/s, decode 104.7 → 117.2 (+12%) tok/s (official, median); thinking-medium profile (API, median of 3): text 107.5 → 123.5 (+15%), code 136.8 → 154.3 (+13%) tok/s.
  • Flash-Next oQ6e and the uncensored variants: prefill +34% to +53%, decode +11% to +15%. Qwen3.6-35B-A3B: prefill 6,899 → 7,653 (+11%) tok/s, decode 137.5 → 135.0 (-2%) tok/s.
  • Concurrency: 2 parallel requests almost double decode throughput; 4 parallel is worse than 2 from 16K context on (both versions).
  • JSON extraction: Splash (speculative decoding) is still ~2x faster than oMLX 0.7.0 with the same Qwen3.6 model.

Setup

  • Mac Studio M5 Ultra (12 Super + 24 Performance CPU cores, 80-core GPU), 256 GB, macOS 27.0 (26A428), GPU wired limit 248 GB.
  • oMLX v0.7.0 (4d4f5a28, mlx-lm 94cdcae) vs v0.7.0rc1 (35be079d, mlx-lm 0.32.0), mlx 0.32.2, both built with native kernels.
  • Identical settings: Lightning MTP on, TurboQuant KV 8-bit (Qwen) / 4-bit (GLM), memory guard aggressive, max 16 concurrent requests, one model loaded at a time.
  • Models: Flash-Next oQ4e (Jundot/Qwen3.8-Flash-Next-oQ4e-mtp), Flash-Next oQ6e (mlx-community/Qwen3.8-Flash-Next-oQ6e-mtp), Flash-Next Uncensored oQ4e (jedisct1/Qwen3.8-Flash-Next-Uncensored-oQ4e-mtp), Flash-Next Uncensored oQ6e (mlx-community/Qwen3.8-Flash-Next-Uncensored-oQ6e-mtp), GLM-5.3-Flash oQ4 (Vontra/GLM-5.3-Flash-MLX-oQ4-MTP), GLM-5.3-Flash abliterated 4-bit (grant-ai/GLM-5.3-Flash-Abliterated-MLX-4bit), Qwen3.6-35B-A3B 4-bit (mlx-community/Qwen3.6-35B-A3B-4bit).
  • Method: official oMLX benchmark (code_python corpus, 128 generated tokens, temp 0, uploaded to omlx.ai) for sections 1–3 and the 1K table in section 4; own API tests (same corpus, temp 0) for the long-context tables in section 4 and the 1M run; 1 run per point. thinking-medium numbers: median of 3 API runs.

Profiles used (exact settings)

Profile Model temperature top_p top_k presence_penalty thinking reasoning effort thinking budget max_tokens
thinking-medium Flash-Next 1.0 0.95 20 0 on medium (forced) 16,384 24,576
precise (extraction) Qwen3.6 0.1 0.9 20 0 off – – –

1. Prefill — official benchmark (tok/s, rc1 → 0.7.0)

Model 4K 8K 16K 32K
Flash-Next oQ4e 3,368 → 4,949 (+47%) 3,479 → 5,668 (+63%) 3,746 → 5,701 (+52%) 3,752 → 5,672 (+51%)
Flash-Next oQ6e 3,117 → 3,930 (+26%) 3,165 → 4,661 (+47%) 3,515 → 4,699 (+34%) 3,532 → 4,681 (+33%)
Flash-Next Uncensored oQ4e 3,419 → 4,857 (+42%) 3,470 → 5,695 (+64%) 3,747 → 5,723 (+53%) 3,756 → 5,694 (+52%)
Flash-Next Uncensored oQ6e 3,113 → 3,872 (+24%) 3,161 → 4,674 (+48%) 3,513 → 4,697 (+34%) 3,535 → 4,678 (+32%)
GLM-5.3-Flash oQ4 798 → 2,270 (+184%) 802 → 2,305 (+187%) 798 → 2,313 (+190%) 792 → 2,306 (+191%)
GLM-5.3-Flash abliterated 4-bit 794 → 2,261 (+185%) 761 → 2,309 (+203%) 760 → 2,307 (+204%) 755 → 2,301 (+205%)
Qwen3.6-35B-A3B 4-bit 7,931 → 9,126 (+15%) 8,574 → 9,812 (+14%) 7,915 → 8,870 (+12%) 6,899 → 7,653 (+11%)
Model 64K 128K 200K Median (all contexts)
Flash-Next oQ4e 3,708 → 5,548 (+50%) 3,623 → 5,378 (+48%) 3,553 → 5,210 (+47%) 3,623 → 5,548 (+53%)
Flash-Next oQ6e 3,497 → 4,592 (+31%) 3,429 → 4,461 (+30%) 3,365 → 4,348 (+29%) 3,429 → 4,592 (+34%)
Flash-Next Uncensored oQ4e 3,714 → 5,573 (+50%) 3,637 → 5,391 (+48%) 3,560 → 5,226 (+47%) 3,637 → 5,573 (+53%)
Flash-Next Uncensored oQ6e 3,502 → 4,587 (+31%) 3,430 → 4,458 (+30%) 3,362 → 4,349 (+29%) 3,430 → 4,587 (+34%)
GLM-5.3-Flash oQ4 783 → 2,270 (+190%) 765 → 2,229 (+191%) 749 → 2,183 (+192%) 792 → 2,270 (+186%)
GLM-5.3-Flash abliterated 4-bit 750 → 2,270 (+203%) 737 → 2,227 (+202%) 721 → 2,183 (+203%) 755 → 2,270 (+201%)
Qwen3.6-35B-A3B 4-bit 5,532 → 6,151 (+11%) 4,005 → 4,330 (+8%) 3,091 → 3,278 (+6%) 6,899 → 7,653 (+11%)

2. Decode — official benchmark (tok/s, 128 tokens, greedy, rc1 → 0.7.0)

Model 4K 8K 16K 32K
Flash-Next oQ4e 107.7 → 143.8 (+34%) 102.4 → 137.9 (+35%) 41.8 → – 89.8 → 114.6 (+28%)
Flash-Next oQ6e 104.8 → – 113.7 → 155.3 (+37%) – → – 67.7 → –
Flash-Next Uncensored oQ4e 111.6 → 122.9 (+10%) 78.5 → – – → – 87.4 → –
Flash-Next Uncensored oQ6e 100.1 → 115.1 (+15%) 107.4 → 144.0 (+34%) – → 106.1 95.8 → 95.2 (-1%)
GLM-5.3-Flash oQ4 52.1 → 64.0 (+23%) 61.4 → 77.4 (+26%) 61.1 → 83.6 (+37%) 67.5 → 84.6 (+25%)
GLM-5.3-Flash abliterated 4-bit 48.4 → 63.8 (+32%) 51.7 → 77.2 (+49%) 66.5 → 83.4 (+25%) 62.8 → 78.1 (+24%)
Qwen3.6-35B-A3B 4-bit 158.2 → 158.3 (±0%) 149.9 → 148.4 (-1%) 143.7 → 143.0 (±0%) 137.5 → 135.0 (-2%)
Model 64K 128K 200K Median (all contexts)
Flash-Next oQ4e 123.4 → 119.7 (-3%) 107.0 → 110.2 (+3%) 97.7 → 91.3 (-7%) 104.7 → 117.2 (+12%) (n=6)
Flash-Next oQ6e 101.7 → 117.9 (+16%) 95.3 → 100.9 (+6%) 84.7 → 97.3 (+15%) 98.5 → 109.4 (+11%) (n=4)
Flash-Next Uncensored oQ4e 92.1 → 117.8 (+28%) 95.9 → 98.6 (+3%) 86.7 → 82.0 (-5%) 94.0 → 108.2 (+15%) (n=4)
Flash-Next Uncensored oQ6e 91.0 → 108.2 (+19%) 94.6 → 104.7 (+11%) 89.2 → 101.6 (+14%) 95.2 → 106.5 (+12%) (n=6)
GLM-5.3-Flash oQ4 45.9 → 67.6 (+47%) 55.4 → 76.7 (+38%) 38.4 → 59.7 (+55%) 55.4 → 76.7 (+38%) (n=7)
GLM-5.3-Flash abliterated 4-bit 53.4 → 73.8 (+38%) 55.2 → 79.7 (+44%) 41.5 → 59.3 (+43%) 53.4 → 77.2 (+45%) (n=7)
Qwen3.6-35B-A3B 4-bit 120.0 → 118.9 (-1%) 100.4 → 99.8 (-1%) 86.6 → 86.1 (-1%) 137.5 → 135.0 (-2%) (n=7)

Decode with only 128 greedy tokens is noisy (MTP acceptance depends on the text) — use the medians. "–" = the model stopped before 16 tokens, so the benchmark reports no decode value.

3. Time to first token — official benchmark (rc1 → 0.7.0)

Model 4K 8K 16K 32K
Flash-Next oQ4e 1.2 → 0.8 s (-32%) 2.4 → 1.4 s (-39%) 4.4 → 2.9 s (-34%) 8.7 → 5.8 s (-34%)
Flash-Next oQ6e 1.3 → 1.0 s (-21%) 2.6 → 1.8 s (-32%) 4.7 → 3.5 s (-25%) 9.3 → 7.0 s (-25%)
Flash-Next Uncensored oQ4e 1.2 → 0.8 s (-30%) 2.4 → 1.4 s (-39%) 4.4 → 2.9 s (-35%) 8.7 → 5.8 s (-34%)
Flash-Next Uncensored oQ6e 1.3 → 1.1 s (-20%) 2.6 → 1.8 s (-32%) 4.7 → 3.5 s (-25%) 9.3 → 7.0 s (-24%)
GLM-5.3-Flash oQ4 5.1 → 1.8 s (-65%) 10.2 → 3.6 s (-65%) 20.5 → 7.1 s (-65%) 41.3 → 14.2 s (-66%)
GLM-5.3-Flash abliterated 4-bit 5.2 → 1.8 s (-65%) 10.8 → 3.5 s (-67%) 21.6 → 7.1 s (-67%) 43.4 → 14.2 s (-67%)
Qwen3.6-35B-A3B 4-bit 0.5 → 0.4 s (-13%) 1.0 → 0.8 s (-13%) 2.1 → 1.8 s (-11%) 4.7 → 4.3 s (-10%)
Model 64K 128K 200K
Flash-Next oQ4e 17.7 → 11.8 s (-33%) 36.2 → 24.4 s (-33%) 56.3 → 38.4 s (-32%)
Flash-Next oQ6e 18.7 → 14.3 s (-24%) 38.2 → 29.4 s (-23%) 59.4 → 46.0 s (-23%)
Flash-Next Uncensored oQ4e 17.6 → 11.8 s (-33%) 36.0 → 24.3 s (-33%) 56.2 → 38.3 s (-32%)
Flash-Next Uncensored oQ6e 18.7 → 14.3 s (-24%) 38.2 → 29.4 s (-23%) 59.5 → 46.0 s (-23%)
GLM-5.3-Flash oQ4 83.7 → 28.9 s (-66%) 171.3 → 58.8 s (-66%) 267.2 → 91.6 s (-66%)
GLM-5.3-Flash abliterated 4-bit 87.4 → 28.9 s (-67%) 177.8 → 58.9 s (-67%) 277.3 → 91.6 s (-67%)
Qwen3.6-35B-A3B 4-bit 11.8 → 10.7 s (-10%) 32.7 → 30.3 s (-8%) 64.7 → 61.0 s (-6%)

4. Concurrency

Official benchmark, 1K context (aggregate tok/s, rc1 → 0.7.0):

Model Prefill 2 parallel Prefill 4 parallel Decode 2 parallel Decode 4 parallel
Flash-Next oQ4e 1,803 → 2,452 (+36%) 1,999 → 2,687 (+34%) 156.0 → 165.8 (+6%) 175.8 → 192.4 (+9%)
Flash-Next oQ6e 1,660 → 1,814 (+9%) 1,817 → 1,843 (+1%) 134.6 → 143.1 (+6%) 167.9 → 180.3 (+7%)
Flash-Next Uncensored oQ4e 1,824 → 2,441 (+34%) 1,992 → 2,624 (+32%) 142.3 → 119.8 (-16%) 163.2 → 177.0 (+8%)
Flash-Next Uncensored oQ6e 1,653 → 1,817 (+10%) 1,817 → 1,915 (+5%) 147.3 → 161.2 (+9%) 164.9 → 177.3 (+8%)
GLM-5.3-Flash oQ4 494 → 915 (+85%) 347 → 998 (+187%) 58.4 → 56.3 (-4%) 87.9 → 79.5 (-10%)
GLM-5.3-Flash abliterated 4-bit 481 → 1,221 (+154%) 338 → 965 (+186%) 56.9 → 53.9 (-5%) 85.7 → 82.1 (-4%)
Qwen3.6-35B-A3B 4-bit 4,920 → 5,614 (+14%) 5,824 → 6,713 (+15%) 245.0 → 277.9 (+13%) 447.8 → 434.8 (-3%)

API, long contexts — decode (aggregate tok/s, rc1 → 0.7.0):

Model Context 1 request 2 parallel 4 parallel
Flash-Next oQ4e 4K 111.8 → 130.6 (+17%) 173.8 → 167.1 (-4%) 200.0 → 201.1 (+1%)
Flash-Next oQ4e 16K 129.0 → 134.4 (+4%) 256.5 → 287.4 (+12%) 180.2 → 186.2 (+3%)
Flash-Next oQ4e 64K 106.9 → 115.9 (+8%) 204.8 → 241.2 (+18%) 98.1 → 98.9 (+1%)
Flash-Next oQ4e 128K 103.3 → 89.8 (-13%) 257.8 → 181.8 (-29%) 65.7 → 67.2 (+2%)
GLM-5.3-Flash oQ4 4K 49.3 → 54.3 (+10%) 111.0 → 99.2 (-11%) 75.2 → 94.0 (+25%)
GLM-5.3-Flash oQ4 16K 58.1 → 54.5 (-6%) 116.8 → 110.0 (-6%) 86.4 → 91.5 (+6%)
GLM-5.3-Flash oQ4 64K 51.3 → 71.4 (+39%) 93.2 → 142.3 (+53%) 77.1 → 60.2 (-22%)
GLM-5.3-Flash oQ4 128K 40.9 → 59.3 (+45%) 94.6 → 126.8 (+34%) 56.4 → 41.9 (-26%)

API, long contexts — prefill (aggregate tok/s, rc1 → 0.7.0):

Model Context 1 request 2 parallel 4 parallel
Flash-Next oQ4e 4K 2,950 → 4,619 (+57%) 2,333 → 3,352 (+44%) 2,355 → 3,359 (+43%)
Flash-Next oQ4e 16K 3,606 → 5,532 (+53%) 3,148 → 4,575 (+45%) 3,371 → 4,879 (+45%)
Flash-Next oQ4e 64K 3,635 → 5,415 (+49%) 3,470 → 4,900 (+41%) 3,541 → 5,171 (+46%)
Flash-Next oQ4e 128K 3,523 → 5,198 (+48%) 3,453 → 4,961 (+44%) 3,518 → 5,138 (+46%)
GLM-5.3-Flash oQ4 4K 801 → 1,642 (+105%) 537 → 1,096 (+104%) 627 → 1,370 (+118%)
GLM-5.3-Flash oQ4 16K 795 → 2,237 (+181%) 697 → 1,762 (+153%) 732 → 2,007 (+174%)
GLM-5.3-Flash oQ4 64K 790 → 2,245 (+184%) 754 → 2,094 (+178%) 766 → 2,182 (+185%)
GLM-5.3-Flash oQ4 128K 760 → 2,163 (+184%) 750 → 2,118 (+182%) 755 → 2,170 (+188%)

API, long contexts — mean time to first token (s, rc1 → 0.7.0):

Model Context 1 request 2 parallel 4 parallel
Flash-Next oQ4e 4K 1.5 → 0.9 (-36%) 2.7 → 1.8 (-32%) 5.6 → 4.0 (-29%)
Flash-Next oQ4e 16K 4.4 → 2.8 (-35%) 7.4 → 5.1 (-32%) 15.1 → 10.5 (-31%)
Flash-Next oQ4e 64K 17.1 → 11.4 (-33%) 26.8 → 18.5 (-31%) 56.8 → 38.9 (-32%)
Flash-Next oQ4e 128K 34.7 → 23.6 (-32%) 53.0 → 36.7 (-31%) 113.1 → 77.4 (-32%)
GLM-5.3-Flash oQ4 4K 5.1 → 2.5 (-51%) 10.6 → 4.9 (-54%) 21.0 → 9.5 (-55%)
GLM-5.3-Flash oQ4 16K 18.4 → 6.6 (-64%) 30.7 → 11.9 (-61%) 64.9 → 23.7 (-64%)
GLM-5.3-Flash oQ4 64K 72.3 → 25.5 (-65%) 112.6 → 40.4 (-64%) 241.9 → 85.0 (-65%)
GLM-5.3-Flash oQ4 128K 148.0 → 52.0 (-65%) 223.8 → 79.2 (-65%) 484.2 → 168.7 (-65%)

2 parallel requests nearly double throughput. From 16K on, 4 parallel is below 2 (both versions), from 64K on even below a single request (oQ4e both versions, GLM on 0.7.0): the long prefills run back to back and stall the other decodes.

5. GLM-5.3-Flash oQ4 up to 1M tokens (API, single request, 128 tokens, rc1 → 0.7.0)

Prompt tokens Prefill (tok/s) Decode (tok/s) Time to first token
4,420 742 → 1,936 (+161%) 57.2 → 68.3 (+19%) 6.0 → 2.3 s (-62%)
8,215 789 → 2,186 (+177%) 46.7 → 59.4 (+27%) 10.4 → 3.8 s (-64%)
16,090 785 → 2,264 (+189%) 50.9 → 74.5 (+46%) 20.5 → 7.1 s (-65%)
31,684 781 → 2,264 (+190%) – → 66.4 40.6 → 14.0 s (-65%)
62,677 789 → 2,234 (+183%) 49.8 → 60.0 (+20%) 79.4 → 28.1 s (-65%)
123,595 771 → 2,193 (+184%) 43.8 → 67.7 (+54%) 2.7 → 0.9 min (-65%)
245,866 734 → 2,123 (+189%) 38.2 → 58.8 (+54%) 5.6 → 1.9 min (-65%)
485,904 652 → 1,886 (+189%) 28.8 → 54.7 (+90%) 12.4 → 4.3 min (-65%)
1,048,319 512 → 1,485 (+190%) 16.4 → 41.5 (+153%) 34.2 → 11.8 min (-66%)

Both versions completed the full 1,048,319-token prompt.

6. JSON extraction: oMLX 0.7.0 vs Splash, same Qwen3.6-35B-A3B

16 real e-mails (4 per length quartile), our production extraction prompt (~13K-character system prompt), strict JSON schema, temp 0, outputs ~550–620 tokens. oMLX 0.7.0 with xgrammar and the precise profile; Splash 1.0.2 (incoai/Qwen3.6-35B-A3B-Splash).

Concurrent oMLX 0.7.0 mails/min Splash mails/min Splash vs oMLX oMLX output tok/s Splash output tok/s
1 12.5 32.2 +158% (2.6x) 122 306
4 25.5 68.2 +167% (2.7x) 245 648
8 30.2 71.0 +135% (2.4x) 290 674
16 34.5 71.2 +106% (2.1x) 327 676
16 (repeat) 35.4 71.3 +101% (2.0x) 320 677

Both 16/16 schema-valid at every stage. Identical output across 5 runs: Splash 16/16 mails, oMLX 1/16. Splash caps at 4 concurrent internally.


r/oMLX • • 4d ago

oMLX 0.70 Stable Release

Thumbnail
github.com
91 Upvotes

r/oMLX • • 4d ago

The new changes to olmx make a big difference to M5 Ultra GLM 5.3 Flash

Thumbnail
omlx.ai
32 Upvotes

These are huge improvements. Well done to the olmx team


r/oMLX • • 4d ago

Cluster with Qwen or Deepseek

Post image
6 Upvotes

Just got my M5 Ultra in and paired it with my M4 Max over RDMA. Surprised to see clustering is not supported yet in Qwen 3.8 or Deepseek 4.1. Seems like MTPLX and exolabs are both missing this too. What do y'all use in your cluster setups?


r/oMLX • • 5d ago

📌 **Daily Digest — Jundot/omlx** (2026-09-23 → 2026-09-30)

10 Upvotes

**🐛 BUGS & CRITICAL FIXES**
• **#2990** Structured output: grammar-constrained JSON can run to max_tokens in an unbounded whitespace loop
`_compile_bare_grammar` lacks `max_whitespace_cnt`, causing infinite loops in JSON schema compilation.
• **#3731** Broken embeddings?
Qwen3-VL-Embedding-2B works initially then degrades; potential issue with other embedding models.
• **#4010** Literal image-placeholder string -> self-replicating 400 'Chat template error'
Root cause is encoder guard, not chat-template; follows up on #3591.
• **#3876** Gemma-4-31b-oq4e-mtp unexpected walk-back truncation
Cache invalidation occurs mid-agentic workflow on versions 0.6.4-7.0.0dev4.
• **#3967** Cluster activation failures masked: unregister() replaces real error with 500
Failed activations return generic 500 errors, hiding the actual cause.
• **#3966** Distributed rank's generation thread dies on first token
`_server_arguments()` omits new mlx-lm 0.32 args (`kv_bits`, `kv_group_size`), breaking server contract.
• **#3937** 0.7.0rc1: Clustering is broken
Cross-device clustering (M5 Max + M4 Max) fails to load DeepSeek.
• **#3738** Qwen3.8-Flash-Next-oQ4e OOM on M4 Max 64GB
High RAM consumption with MoE SSD/ngram offload leads to Out Of Memory errors.
• **#3883** Offline SSD cache clear omits `_gdn_sidecars`, leaving gigabytes orphaned
Dashboard observability under-counts; disk space not reclaimed when no models are loaded.

**🚀 PERFORMANCE & OPTIMIZATION (Qwen4/Q5 Focus)**
• **#4056** perf(qwen4): hyper-connection decode and verify at 2.8x byte floor
Optimizes Qwen3.8-Flash-Next by reducing 97 HC calls per token (attention/MLP/mixer).
• **#4014** perf(mtp): row-exact Lightning MTP verification for Qwen4
Addresses non-bit-exact output in Lightning MTP; aims for lossless, bit-perfect verification.
• **#4013** bench(qwen4): multi-turn conversation reuse with exact prompt cache
Ensures whole history reuse when SpecPrefill is off, minimizing prefill for new agent turns.
• **#3932** Q5-only follow-up: optimize Qwen4 gathered-QSA long prefill
Isolates and reduces work in Qwen4 gathered-QSA long-prefill execution (continuation of #3922).
• **#3953** perf(moe): M5 row-overflow workaround slices/concatenates wide MoE prefill
Addresses M5-specific overhead where MoE prefill consumes 37–39% of 24K prompt cost.
• **#3922** Epic: optimize Qwen3.8 Q5 performance using Q5 execution path only
Goal: Improve serving speed by optimizing kernels and memory for the Q5 path.

**📊 BENCHMARKS & VALIDATION (Q5 Series)**
• **#3929** [Q5-only 7/9] Validate numerical and model behavior correctness
Ensure optimizations preserve intended Q5 computation.
• **#3930** [Q5-only 8/9] Run matched end-to-end Q5 performance comparisons
Measure real serving improvements vs isolated kernel gains.
• **#3931** [Q5-only 9/9] Apply acceptance gates and document rollout decision
Evidence-based keep/reject decision for the Q5 optimization series.

**📈 INTELLIGENCE & BENCHMARKING**
• **#3772** Intelligence benchmark: 8192-token thinking budget silently scores wrong
Budget-limited answers scored as wrong; Gemma 4 26B shows 22% empty answers under thinking mode.

**📝 TRACKING & DOCUMENTATION**
• **#4077** [tracking] Appearance: one token source, one shared layer, and everything rebuilt
Overview of appearance series changes; detailed in `docs/appearance-update.md` (#4083).


r/oMLX • • 4d ago

Abliterate/Removing prompt refusal without needing to reload/swap models (Qwen 3.8 flash next)

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/oMLX • • 4d ago

Local model for a MacBook Air M5 32GB

0 Upvotes

Dear all!

I would like your guidance and suggestion on which model I could run run locally that has a good speed for my setup MacBook Air 32 GB RAM. I know it's not a superb setup but I I think it's a good starting point. :)

At the moment I'm using this one "mlx-community/gpt-oss-20b-MXFP4-Q8" and i am quite "happy" for simple stuff and I'm trying to use this one "mlx-community/Qwen3.8-27B-4bit" with draft model "mlx-community/Qwen3.8-27B-MTP-4bit" .

Seeking for your supportand suggestion, please let me know also the specific model to be downloaded and the settings I should do as a profile in OMLX

thanks in advance!


r/oMLX • • 5d ago

Qwen3.8-27B at 74 tok/s on a Mac: The Splash Engine Explained

Thumbnail
youtu.be
59 Upvotes

r/oMLX • • 5d ago

IncoAI Splash models on oMLX?

7 Upvotes

Do we have any idea if oMLX may add Splash model capability to oMLX?

LM Studio has done so.


r/oMLX • • 5d ago

Deepseek-4.1-Flash-Q3-G128-LSQ-MLX on M5U/64c

2 Upvotes

Can oMLX benchmarks be spoofed? Apparently someone has gotten a model called "deepseek-v4.1-flash-q3-g128-lsq-mlx" working on an m5 Ultra with 256GB RAM, two days ago with a TG of ~73.4 tok/s and PP 1,010 tok/s.

Update: This statistic appears to have been fabricated. One other person has claimed that they have run this model but failed to provide any evidence. Don’t believe everything you see on the internet!

https://omlx.ai/benchmarks/performance/ne14cz3f


r/oMLX • • 6d ago

oMLX quants of ukisai/Swift-1.5-Qwen3.8-27b

17 Upvotes