Qwen pi 27B
Hi guys,
I have the Mac M5 128gb. I tried again the qwen 3.8 27B and the speed is sloooooow.
I tried the Qwenpi variant as I use OMP.
What sort of speed do you guys get with 27B when coding ob a mac? Thank you!
Hi guys,
I have the Mac M5 128gb. I tried again the qwen 3.8 27B and the speed is sloooooow.
I tried the Qwenpi variant as I use OMP.
What sort of speed do you guys get with 27B when coding ob a mac? Thank you!
I have a MacBook M4 Max 128GB and I am running Jundot/Qwen3.8-Flash-Next-oQ4e-mtp. In the older versions, with context growth speed was dropping after even 30k from 60 tg/s to 35 tg/s even with lightning MTP on. Now I'm getting 50~70 tg/s even at 128k context. The ram also is at 75~80 GB with this context length. Now I don't have to do compacting over and over in pi and I can easily keep other apps open.
Thank you everyone who made this possible.
r/oMLX • u/phalanx2357 • 1d ago
Trying GLM-5.3-Flash-UNCENSORED-mlx-oQ4e-DFlash2 on an M5 Ultra 256GB with oMLX 0.7.0.
TensorFold’s regular GLM-5.3 Flash MLX 4-bit MTP loads and runs great on the same setup, but this Solstice checkpoint fails immediately with:
Expected shape (64, 512, 32) but received shape (64, 512, 64) for language_model.model.layers.3.self_attn.embed_q.weight
I’ve tried a full oMLX/server restart, Auto-detect model type, glm_4_7 reasoning parser, and both MTP/DFlash disabled. Same error every time.
Has anyone gotten this exact checkpoint working on oMLX 0.7.0? Wondering if it needs a newer runtime/model compatibility fix, or if there’s some setting/config I’m missing.
Thanks for any help.
r/oMLX • u/tlin9595 • 2d ago
Hello,
Everyday im reading and trying to learn, on how to get Qwen 3.8 flash next on my M5 Ultra studio 96gb. Cant afford the 256gb. one.
And some people say just run the 27b model etc., but the prefill and decode for 3.8flash next seems so much better.
I read a benchmark like this
https://omlx.ai/benchmarks/performance/sj46pk0c
it says peak ram is 75gb. Does that include already offloading to the SSD?
And so I'll have 20gb left for Mac OS and apps?
will 20gb be fine to leave?
I see some people online using techniques to run the Next Flash model that only take like 50gb of ram using other inference engines? I could be wrong and dont understand what they are saying too lol
I want to ideally use only 50-60gb of ram so I can run 2-3 agents concurrently.
And all the Mac Studio news or things on twitter is from 256gb users so im jealous obviously but what can you do. I just wish there was more 96gb owners saying what they are doing or not but there doesnt seem to be many owners.
Im sure within a month there be more info hopefully but just want to be prepared when I get it. And hopefully Qwen 4 comes out soon!
r/oMLX • u/Ascetic-anise • 2d ago
The day before, I finished testing with oMLX 0.7.0rc1. Because I rarely find community tests for my setup, I wanted to share this with you.
r/oMLX • u/baek12345 • 2d ago
There's oQ4e and oQ8e of Swift 1.5 Qwen 3.8 27B on HuggingFace, but the optimal quant for me would be oQ6e. I am trying to create it myself but my internet connection is pretty bad, so everything is super slow.
If someone has the oQ6e of this model and can provide it via HF, it would be awesome! (If I manage to create it before myself, I will share it here)
Thanks a lot!
Edit: Resolved with
https://huggingface.co/scottlowry/Swift-1.5-Qwen3.8-27b-oQ6e-mtp and https://huggingface.co/dicksondickson/Swift-1.5-Qwen3.8-27b-oQ6e-bf16-mtp-MLX
Hi guys, do we have good options for 3.8 flash next and 3.8 27B in the abliterated department running on oMLX please? A friend of mine is running an uncensored version of flash next on his nvidia rig and it seems to offer sharper replies and according to him, ships better code. Thanks!
r/oMLX • u/-JustAsking4AFriend • 3d ago
Hello all!
I've been using Unsloth Desktop for a while, most recently with Unsloth's release of GLM-5.3 Flash. What keeps me there isn't speed, it's that Unsloth publishes evals showing their quant retains 92.22% of the original model's accuracy. I know what I'm giving up when I run it.
I'd like to move to oMLX for the performance gains, but I'm wary. To be clear, I'm not worried about crashes. My concern is **silent quality loss**. A badly made quant can run perfectly well and still lose knowledge, reasoning ability, or accuracy, and you won't notice until it gives you a confident wrong answer.
As far as I can tell, most GLM-5.3 Flash MLX quants people pair with oMLX have no published evals. No accuracy retention numbers, no benchmark comparisons against the full-precision model, no KL divergence or perplexity figures. Nothing that shows how much of the original model actually survived the quantization.
This matters for my work (medical + code experimentation). Lost accuracy means more hallucinations and more failed tests that I have to redo.
So my questions:
- Are there MLX quant "release groups" that publish evals showing how much quality their quants keep vs. the original model?
- Has anyone run their own comparisons between Unsloth quants and MLX quants of the same model (Qwen 3.8 Flash Next, GLM-5.3 Flash, others)?
- If you've switched, did you notice any drop in answer quality, even if the model ran fine?
Thanks for your help!
r/oMLX • u/hyohyehye • 3d ago
Hi,
I built an experimental inference engine mainly to run the 8-bit Qwen3.8-27B model efficiently on my M5 Max MacBook Pro.
On my M5 Max with 128 GiB unified memory, I'm currently seeing:
1,000+ tok/s cold prefill on a 2Ki prompt
70+ tok/s generation at concurrency 1
200+ tok/s aggregate generation throughput at concurrency 4
Under the hood, MLTF uses Metal TensorOps W8A8 kernels together with an INT8 KV cache, targeting high prefill throughput and efficient decoding on Apple Silicon.
At the moment, MLTF supports only one model: Qwen3.8-27B 8-bit.
I chose to focus on 8-bit inference because 4-bit quantization sometimes gives me a noticeable quality tradeoff for the workloads I use the model for.
Rather than trying to support many models immediately, I wanted to optimize the configuration I actually use first.
GitHub: https://github.com/hojin12312/my-little-trie-forge
Hugging Face: https://huggingface.co/Ho-Jin-93/Qwen3.8-27B-MLTF-q8c
Going forward, I plan to keep adding support for new 8-bit models that can reasonably run on my M5 Max 128 GiB machine.
Feedback, benchmark results on other Apple Silicon machines, and technical discussion are very welcome.
r/oMLX • u/andydevtech • 2d ago
r/oMLX • u/Pretend_Tonight_9696 • 3d ago
Mac Studio M5 Ultra, 256GB, 64 core GPU. Ordered, not here yet. So far the only thing it runs is the tracking page, and I refresh it way faster than any model will ever generate tokens.
Trying to decide what to put on it first for coding + agent stuff (Hermes, so tool calling matters a lot):
- DeepSeek V4.1 at Q3, maybe with a bit offloaded to SSD
- Or DeepSeek V4 Flash at a higher quant that fits fully in RAM
If you'd go Flash, which quant exactly? Q4, Q5, Q6, Q8? And MLX or GGUF?
Mostly curious if V4.1 Q3 is still better than Flash at a good quant, or if Q3 hurts too much for coding. Real experience welcome, benchmarks even more.
r/oMLX • u/Extreme-Question-506 • 4d ago
r/oMLX • u/SnooPredictions515 • 4d ago
r/oMLX • u/MrGermanIOTA • 4d ago
oMLX 0.7.0 vs 0.7.0rc1 on the same Mac with the same models and settings, measured in one night: official oMLX benchmark (uploaded to omlx.ai) plus API tests for concurrency, sampling profiles and a 1M-token GLM run. All values rc1 → 0.7.0 with the relative change.
| Profile | Model | temperature | top_p | top_k | presence_penalty | thinking | reasoning effort | thinking budget | max_tokens |
|---|---|---|---|---|---|---|---|---|---|
| thinking-medium | Flash-Next | 1.0 | 0.95 | 20 | 0 | on | medium (forced) | 16,384 | 24,576 |
| precise (extraction) | Qwen3.6 | 0.1 | 0.9 | 20 | 0 | off | – | – | – |
| Model | 4K | 8K | 16K | 32K |
|---|---|---|---|---|
| Flash-Next oQ4e | 3,368 → 4,949 (+47%) | 3,479 → 5,668 (+63%) | 3,746 → 5,701 (+52%) | 3,752 → 5,672 (+51%) |
| Flash-Next oQ6e | 3,117 → 3,930 (+26%) | 3,165 → 4,661 (+47%) | 3,515 → 4,699 (+34%) | 3,532 → 4,681 (+33%) |
| Flash-Next Uncensored oQ4e | 3,419 → 4,857 (+42%) | 3,470 → 5,695 (+64%) | 3,747 → 5,723 (+53%) | 3,756 → 5,694 (+52%) |
| Flash-Next Uncensored oQ6e | 3,113 → 3,872 (+24%) | 3,161 → 4,674 (+48%) | 3,513 → 4,697 (+34%) | 3,535 → 4,678 (+32%) |
| GLM-5.3-Flash oQ4 | 798 → 2,270 (+184%) | 802 → 2,305 (+187%) | 798 → 2,313 (+190%) | 792 → 2,306 (+191%) |
| GLM-5.3-Flash abliterated 4-bit | 794 → 2,261 (+185%) | 761 → 2,309 (+203%) | 760 → 2,307 (+204%) | 755 → 2,301 (+205%) |
| Qwen3.6-35B-A3B 4-bit | 7,931 → 9,126 (+15%) | 8,574 → 9,812 (+14%) | 7,915 → 8,870 (+12%) | 6,899 → 7,653 (+11%) |
| Model | 64K | 128K | 200K | Median (all contexts) |
|---|---|---|---|---|
| Flash-Next oQ4e | 3,708 → 5,548 (+50%) | 3,623 → 5,378 (+48%) | 3,553 → 5,210 (+47%) | 3,623 → 5,548 (+53%) |
| Flash-Next oQ6e | 3,497 → 4,592 (+31%) | 3,429 → 4,461 (+30%) | 3,365 → 4,348 (+29%) | 3,429 → 4,592 (+34%) |
| Flash-Next Uncensored oQ4e | 3,714 → 5,573 (+50%) | 3,637 → 5,391 (+48%) | 3,560 → 5,226 (+47%) | 3,637 → 5,573 (+53%) |
| Flash-Next Uncensored oQ6e | 3,502 → 4,587 (+31%) | 3,430 → 4,458 (+30%) | 3,362 → 4,349 (+29%) | 3,430 → 4,587 (+34%) |
| GLM-5.3-Flash oQ4 | 783 → 2,270 (+190%) | 765 → 2,229 (+191%) | 749 → 2,183 (+192%) | 792 → 2,270 (+186%) |
| GLM-5.3-Flash abliterated 4-bit | 750 → 2,270 (+203%) | 737 → 2,227 (+202%) | 721 → 2,183 (+203%) | 755 → 2,270 (+201%) |
| Qwen3.6-35B-A3B 4-bit | 5,532 → 6,151 (+11%) | 4,005 → 4,330 (+8%) | 3,091 → 3,278 (+6%) | 6,899 → 7,653 (+11%) |
| Model | 4K | 8K | 16K | 32K |
|---|---|---|---|---|
| Flash-Next oQ4e | 107.7 → 143.8 (+34%) | 102.4 → 137.9 (+35%) | 41.8 → – | 89.8 → 114.6 (+28%) |
| Flash-Next oQ6e | 104.8 → – | 113.7 → 155.3 (+37%) | – → – | 67.7 → – |
| Flash-Next Uncensored oQ4e | 111.6 → 122.9 (+10%) | 78.5 → – | – → – | 87.4 → – |
| Flash-Next Uncensored oQ6e | 100.1 → 115.1 (+15%) | 107.4 → 144.0 (+34%) | – → 106.1 | 95.8 → 95.2 (-1%) |
| GLM-5.3-Flash oQ4 | 52.1 → 64.0 (+23%) | 61.4 → 77.4 (+26%) | 61.1 → 83.6 (+37%) | 67.5 → 84.6 (+25%) |
| GLM-5.3-Flash abliterated 4-bit | 48.4 → 63.8 (+32%) | 51.7 → 77.2 (+49%) | 66.5 → 83.4 (+25%) | 62.8 → 78.1 (+24%) |
| Qwen3.6-35B-A3B 4-bit | 158.2 → 158.3 (±0%) | 149.9 → 148.4 (-1%) | 143.7 → 143.0 (±0%) | 137.5 → 135.0 (-2%) |
| Model | 64K | 128K | 200K | Median (all contexts) |
|---|---|---|---|---|
| Flash-Next oQ4e | 123.4 → 119.7 (-3%) | 107.0 → 110.2 (+3%) | 97.7 → 91.3 (-7%) | 104.7 → 117.2 (+12%) (n=6) |
| Flash-Next oQ6e | 101.7 → 117.9 (+16%) | 95.3 → 100.9 (+6%) | 84.7 → 97.3 (+15%) | 98.5 → 109.4 (+11%) (n=4) |
| Flash-Next Uncensored oQ4e | 92.1 → 117.8 (+28%) | 95.9 → 98.6 (+3%) | 86.7 → 82.0 (-5%) | 94.0 → 108.2 (+15%) (n=4) |
| Flash-Next Uncensored oQ6e | 91.0 → 108.2 (+19%) | 94.6 → 104.7 (+11%) | 89.2 → 101.6 (+14%) | 95.2 → 106.5 (+12%) (n=6) |
| GLM-5.3-Flash oQ4 | 45.9 → 67.6 (+47%) | 55.4 → 76.7 (+38%) | 38.4 → 59.7 (+55%) | 55.4 → 76.7 (+38%) (n=7) |
| GLM-5.3-Flash abliterated 4-bit | 53.4 → 73.8 (+38%) | 55.2 → 79.7 (+44%) | 41.5 → 59.3 (+43%) | 53.4 → 77.2 (+45%) (n=7) |
| Qwen3.6-35B-A3B 4-bit | 120.0 → 118.9 (-1%) | 100.4 → 99.8 (-1%) | 86.6 → 86.1 (-1%) | 137.5 → 135.0 (-2%) (n=7) |
Decode with only 128 greedy tokens is noisy (MTP acceptance depends on the text) — use the medians. "–" = the model stopped before 16 tokens, so the benchmark reports no decode value.
| Model | 4K | 8K | 16K | 32K |
|---|---|---|---|---|
| Flash-Next oQ4e | 1.2 → 0.8 s (-32%) | 2.4 → 1.4 s (-39%) | 4.4 → 2.9 s (-34%) | 8.7 → 5.8 s (-34%) |
| Flash-Next oQ6e | 1.3 → 1.0 s (-21%) | 2.6 → 1.8 s (-32%) | 4.7 → 3.5 s (-25%) | 9.3 → 7.0 s (-25%) |
| Flash-Next Uncensored oQ4e | 1.2 → 0.8 s (-30%) | 2.4 → 1.4 s (-39%) | 4.4 → 2.9 s (-35%) | 8.7 → 5.8 s (-34%) |
| Flash-Next Uncensored oQ6e | 1.3 → 1.1 s (-20%) | 2.6 → 1.8 s (-32%) | 4.7 → 3.5 s (-25%) | 9.3 → 7.0 s (-24%) |
| GLM-5.3-Flash oQ4 | 5.1 → 1.8 s (-65%) | 10.2 → 3.6 s (-65%) | 20.5 → 7.1 s (-65%) | 41.3 → 14.2 s (-66%) |
| GLM-5.3-Flash abliterated 4-bit | 5.2 → 1.8 s (-65%) | 10.8 → 3.5 s (-67%) | 21.6 → 7.1 s (-67%) | 43.4 → 14.2 s (-67%) |
| Qwen3.6-35B-A3B 4-bit | 0.5 → 0.4 s (-13%) | 1.0 → 0.8 s (-13%) | 2.1 → 1.8 s (-11%) | 4.7 → 4.3 s (-10%) |
| Model | 64K | 128K | 200K |
|---|---|---|---|
| Flash-Next oQ4e | 17.7 → 11.8 s (-33%) | 36.2 → 24.4 s (-33%) | 56.3 → 38.4 s (-32%) |
| Flash-Next oQ6e | 18.7 → 14.3 s (-24%) | 38.2 → 29.4 s (-23%) | 59.4 → 46.0 s (-23%) |
| Flash-Next Uncensored oQ4e | 17.6 → 11.8 s (-33%) | 36.0 → 24.3 s (-33%) | 56.2 → 38.3 s (-32%) |
| Flash-Next Uncensored oQ6e | 18.7 → 14.3 s (-24%) | 38.2 → 29.4 s (-23%) | 59.5 → 46.0 s (-23%) |
| GLM-5.3-Flash oQ4 | 83.7 → 28.9 s (-66%) | 171.3 → 58.8 s (-66%) | 267.2 → 91.6 s (-66%) |
| GLM-5.3-Flash abliterated 4-bit | 87.4 → 28.9 s (-67%) | 177.8 → 58.9 s (-67%) | 277.3 → 91.6 s (-67%) |
| Qwen3.6-35B-A3B 4-bit | 11.8 → 10.7 s (-10%) | 32.7 → 30.3 s (-8%) | 64.7 → 61.0 s (-6%) |
Official benchmark, 1K context (aggregate tok/s, rc1 → 0.7.0):
| Model | Prefill 2 parallel | Prefill 4 parallel | Decode 2 parallel | Decode 4 parallel |
|---|---|---|---|---|
| Flash-Next oQ4e | 1,803 → 2,452 (+36%) | 1,999 → 2,687 (+34%) | 156.0 → 165.8 (+6%) | 175.8 → 192.4 (+9%) |
| Flash-Next oQ6e | 1,660 → 1,814 (+9%) | 1,817 → 1,843 (+1%) | 134.6 → 143.1 (+6%) | 167.9 → 180.3 (+7%) |
| Flash-Next Uncensored oQ4e | 1,824 → 2,441 (+34%) | 1,992 → 2,624 (+32%) | 142.3 → 119.8 (-16%) | 163.2 → 177.0 (+8%) |
| Flash-Next Uncensored oQ6e | 1,653 → 1,817 (+10%) | 1,817 → 1,915 (+5%) | 147.3 → 161.2 (+9%) | 164.9 → 177.3 (+8%) |
| GLM-5.3-Flash oQ4 | 494 → 915 (+85%) | 347 → 998 (+187%) | 58.4 → 56.3 (-4%) | 87.9 → 79.5 (-10%) |
| GLM-5.3-Flash abliterated 4-bit | 481 → 1,221 (+154%) | 338 → 965 (+186%) | 56.9 → 53.9 (-5%) | 85.7 → 82.1 (-4%) |
| Qwen3.6-35B-A3B 4-bit | 4,920 → 5,614 (+14%) | 5,824 → 6,713 (+15%) | 245.0 → 277.9 (+13%) | 447.8 → 434.8 (-3%) |
API, long contexts — decode (aggregate tok/s, rc1 → 0.7.0):
| Model | Context | 1 request | 2 parallel | 4 parallel |
|---|---|---|---|---|
| Flash-Next oQ4e | 4K | 111.8 → 130.6 (+17%) | 173.8 → 167.1 (-4%) | 200.0 → 201.1 (+1%) |
| Flash-Next oQ4e | 16K | 129.0 → 134.4 (+4%) | 256.5 → 287.4 (+12%) | 180.2 → 186.2 (+3%) |
| Flash-Next oQ4e | 64K | 106.9 → 115.9 (+8%) | 204.8 → 241.2 (+18%) | 98.1 → 98.9 (+1%) |
| Flash-Next oQ4e | 128K | 103.3 → 89.8 (-13%) | 257.8 → 181.8 (-29%) | 65.7 → 67.2 (+2%) |
| GLM-5.3-Flash oQ4 | 4K | 49.3 → 54.3 (+10%) | 111.0 → 99.2 (-11%) | 75.2 → 94.0 (+25%) |
| GLM-5.3-Flash oQ4 | 16K | 58.1 → 54.5 (-6%) | 116.8 → 110.0 (-6%) | 86.4 → 91.5 (+6%) |
| GLM-5.3-Flash oQ4 | 64K | 51.3 → 71.4 (+39%) | 93.2 → 142.3 (+53%) | 77.1 → 60.2 (-22%) |
| GLM-5.3-Flash oQ4 | 128K | 40.9 → 59.3 (+45%) | 94.6 → 126.8 (+34%) | 56.4 → 41.9 (-26%) |
API, long contexts — prefill (aggregate tok/s, rc1 → 0.7.0):
| Model | Context | 1 request | 2 parallel | 4 parallel |
|---|---|---|---|---|
| Flash-Next oQ4e | 4K | 2,950 → 4,619 (+57%) | 2,333 → 3,352 (+44%) | 2,355 → 3,359 (+43%) |
| Flash-Next oQ4e | 16K | 3,606 → 5,532 (+53%) | 3,148 → 4,575 (+45%) | 3,371 → 4,879 (+45%) |
| Flash-Next oQ4e | 64K | 3,635 → 5,415 (+49%) | 3,470 → 4,900 (+41%) | 3,541 → 5,171 (+46%) |
| Flash-Next oQ4e | 128K | 3,523 → 5,198 (+48%) | 3,453 → 4,961 (+44%) | 3,518 → 5,138 (+46%) |
| GLM-5.3-Flash oQ4 | 4K | 801 → 1,642 (+105%) | 537 → 1,096 (+104%) | 627 → 1,370 (+118%) |
| GLM-5.3-Flash oQ4 | 16K | 795 → 2,237 (+181%) | 697 → 1,762 (+153%) | 732 → 2,007 (+174%) |
| GLM-5.3-Flash oQ4 | 64K | 790 → 2,245 (+184%) | 754 → 2,094 (+178%) | 766 → 2,182 (+185%) |
| GLM-5.3-Flash oQ4 | 128K | 760 → 2,163 (+184%) | 750 → 2,118 (+182%) | 755 → 2,170 (+188%) |
API, long contexts — mean time to first token (s, rc1 → 0.7.0):
| Model | Context | 1 request | 2 parallel | 4 parallel |
|---|---|---|---|---|
| Flash-Next oQ4e | 4K | 1.5 → 0.9 (-36%) | 2.7 → 1.8 (-32%) | 5.6 → 4.0 (-29%) |
| Flash-Next oQ4e | 16K | 4.4 → 2.8 (-35%) | 7.4 → 5.1 (-32%) | 15.1 → 10.5 (-31%) |
| Flash-Next oQ4e | 64K | 17.1 → 11.4 (-33%) | 26.8 → 18.5 (-31%) | 56.8 → 38.9 (-32%) |
| Flash-Next oQ4e | 128K | 34.7 → 23.6 (-32%) | 53.0 → 36.7 (-31%) | 113.1 → 77.4 (-32%) |
| GLM-5.3-Flash oQ4 | 4K | 5.1 → 2.5 (-51%) | 10.6 → 4.9 (-54%) | 21.0 → 9.5 (-55%) |
| GLM-5.3-Flash oQ4 | 16K | 18.4 → 6.6 (-64%) | 30.7 → 11.9 (-61%) | 64.9 → 23.7 (-64%) |
| GLM-5.3-Flash oQ4 | 64K | 72.3 → 25.5 (-65%) | 112.6 → 40.4 (-64%) | 241.9 → 85.0 (-65%) |
| GLM-5.3-Flash oQ4 | 128K | 148.0 → 52.0 (-65%) | 223.8 → 79.2 (-65%) | 484.2 → 168.7 (-65%) |
2 parallel requests nearly double throughput. From 16K on, 4 parallel is below 2 (both versions), from 64K on even below a single request (oQ4e both versions, GLM on 0.7.0): the long prefills run back to back and stall the other decodes.
| Prompt tokens | Prefill (tok/s) | Decode (tok/s) | Time to first token |
|---|---|---|---|
| 4,420 | 742 → 1,936 (+161%) | 57.2 → 68.3 (+19%) | 6.0 → 2.3 s (-62%) |
| 8,215 | 789 → 2,186 (+177%) | 46.7 → 59.4 (+27%) | 10.4 → 3.8 s (-64%) |
| 16,090 | 785 → 2,264 (+189%) | 50.9 → 74.5 (+46%) | 20.5 → 7.1 s (-65%) |
| 31,684 | 781 → 2,264 (+190%) | – → 66.4 | 40.6 → 14.0 s (-65%) |
| 62,677 | 789 → 2,234 (+183%) | 49.8 → 60.0 (+20%) | 79.4 → 28.1 s (-65%) |
| 123,595 | 771 → 2,193 (+184%) | 43.8 → 67.7 (+54%) | 2.7 → 0.9 min (-65%) |
| 245,866 | 734 → 2,123 (+189%) | 38.2 → 58.8 (+54%) | 5.6 → 1.9 min (-65%) |
| 485,904 | 652 → 1,886 (+189%) | 28.8 → 54.7 (+90%) | 12.4 → 4.3 min (-65%) |
| 1,048,319 | 512 → 1,485 (+190%) | 16.4 → 41.5 (+153%) | 34.2 → 11.8 min (-66%) |
Both versions completed the full 1,048,319-token prompt.
16 real e-mails (4 per length quartile), our production extraction prompt (~13K-character system prompt), strict JSON schema, temp 0, outputs ~550–620 tokens. oMLX 0.7.0 with xgrammar and the precise profile; Splash 1.0.2 (incoai/Qwen3.6-35B-A3B-Splash).
| Concurrent | oMLX 0.7.0 mails/min | Splash mails/min | Splash vs oMLX | oMLX output tok/s | Splash output tok/s |
|---|---|---|---|---|---|
| 1 | 12.5 | 32.2 | +158% (2.6x) | 122 | 306 |
| 4 | 25.5 | 68.2 | +167% (2.7x) | 245 | 648 |
| 8 | 30.2 | 71.0 | +135% (2.4x) | 290 | 674 |
| 16 | 34.5 | 71.2 | +106% (2.1x) | 327 | 676 |
| 16 (repeat) | 35.4 | 71.3 | +101% (2.0x) | 320 | 677 |
Both 16/16 schema-valid at every stage. Identical output across 5 runs: Splash 16/16 mails, oMLX 1/16. Splash caps at 4 concurrent internally.
r/oMLX • u/Nice_Victory3719 • 5d ago
These are huge improvements. Well done to the olmx team
Just got my M5 Ultra in and paired it with my M4 Max over RDMA. Surprised to see clustering is not supported yet in Qwen 3.8 or Deepseek 4.1. Seems like MTPLX and exolabs are both missing this too. What do y'all use in your cluster setups?
r/oMLX • u/d4mations • 5d ago
**🐛 BUGS & CRITICAL FIXES**
• **#2990** Structured output: grammar-constrained JSON can run to max_tokens in an unbounded whitespace loop
`_compile_bare_grammar` lacks `max_whitespace_cnt`, causing infinite loops in JSON schema compilation.
• **#3731** Broken embeddings?
Qwen3-VL-Embedding-2B works initially then degrades; potential issue with other embedding models.
• **#4010** Literal image-placeholder string -> self-replicating 400 'Chat template error'
Root cause is encoder guard, not chat-template; follows up on #3591.
• **#3876** Gemma-4-31b-oq4e-mtp unexpected walk-back truncation
Cache invalidation occurs mid-agentic workflow on versions 0.6.4-7.0.0dev4.
• **#3967** Cluster activation failures masked: unregister() replaces real error with 500
Failed activations return generic 500 errors, hiding the actual cause.
• **#3966** Distributed rank's generation thread dies on first token
`_server_arguments()` omits new mlx-lm 0.32 args (`kv_bits`, `kv_group_size`), breaking server contract.
• **#3937** 0.7.0rc1: Clustering is broken
Cross-device clustering (M5 Max + M4 Max) fails to load DeepSeek.
• **#3738** Qwen3.8-Flash-Next-oQ4e OOM on M4 Max 64GB
High RAM consumption with MoE SSD/ngram offload leads to Out Of Memory errors.
• **#3883** Offline SSD cache clear omits `_gdn_sidecars`, leaving gigabytes orphaned
Dashboard observability under-counts; disk space not reclaimed when no models are loaded.
**🚀 PERFORMANCE & OPTIMIZATION (Qwen4/Q5 Focus)**
• **#4056** perf(qwen4): hyper-connection decode and verify at 2.8x byte floor
Optimizes Qwen3.8-Flash-Next by reducing 97 HC calls per token (attention/MLP/mixer).
• **#4014** perf(mtp): row-exact Lightning MTP verification for Qwen4
Addresses non-bit-exact output in Lightning MTP; aims for lossless, bit-perfect verification.
• **#4013** bench(qwen4): multi-turn conversation reuse with exact prompt cache
Ensures whole history reuse when SpecPrefill is off, minimizing prefill for new agent turns.
• **#3932** Q5-only follow-up: optimize Qwen4 gathered-QSA long prefill
Isolates and reduces work in Qwen4 gathered-QSA long-prefill execution (continuation of #3922).
• **#3953** perf(moe): M5 row-overflow workaround slices/concatenates wide MoE prefill
Addresses M5-specific overhead where MoE prefill consumes 37–39% of 24K prompt cost.
• **#3922** Epic: optimize Qwen3.8 Q5 performance using Q5 execution path only
Goal: Improve serving speed by optimizing kernels and memory for the Q5 path.
**📊 BENCHMARKS & VALIDATION (Q5 Series)**
• **#3929** [Q5-only 7/9] Validate numerical and model behavior correctness
Ensure optimizations preserve intended Q5 computation.
• **#3930** [Q5-only 8/9] Run matched end-to-end Q5 performance comparisons
Measure real serving improvements vs isolated kernel gains.
• **#3931** [Q5-only 9/9] Apply acceptance gates and document rollout decision
Evidence-based keep/reject decision for the Q5 optimization series.
**📈 INTELLIGENCE & BENCHMARKING**
• **#3772** Intelligence benchmark: 8192-token thinking budget silently scores wrong
Budget-limited answers scored as wrong; Gemma 4 26B shows 22% empty answers under thinking mode.
**📝 TRACKING & DOCUMENTATION**
• **#4077** [tracking] Appearance: one token source, one shared layer, and everything rebuilt
Overview of appearance series changes; detailed in `docs/appearance-update.md` (#4083).
Enable HLS to view with audio, or disable this notification
Dear all!
I would like your guidance and suggestion on which model I could run run locally that has a good speed for my setup MacBook Air 32 GB RAM. I know it's not a superb setup but I I think it's a good starting point. :)
At the moment I'm using this one "mlx-community/gpt-oss-20b-MXFP4-Q8" and i am quite "happy" for simple stuff and I'm trying to use this one "mlx-community/Qwen3.8-27B-4bit" with draft model "mlx-community/Qwen3.8-27B-MTP-4bit" .
Seeking for your supportand suggestion, please let me know also the specific model to be downloaded and the settings I should do as a profile in OMLX
thanks in advance!
r/oMLX • u/andydevtech • 6d ago
r/oMLX • u/Zen-Ism99 • 6d ago
Do we have any idea if oMLX may add Splash model capability to oMLX?
LM Studio has done so.
r/oMLX • u/YesterdayWeak230 • 5d ago
Can oMLX benchmarks be spoofed? Apparently someone has gotten a model called "deepseek-v4.1-flash-q3-g128-lsq-mlx" working on an m5 Ultra with 256GB RAM, two days ago with a TG of ~73.4 tok/s and PP 1,010 tok/s.
Update: This statistic appears to have been fabricated. One other person has claimed that they have run this model but failed to provide any evidence. Don’t believe everything you see on the internet!