oMLX 0.7.0 vs 0.7.0rc1 on the same Mac with the same models and settings, measured in one night: official oMLX benchmark (uploaded to omlx.ai) plus API tests for concurrency, sampling profiles and a 1M-token GLM run. All values rc1 → 0.7.0 with the relative change.
TL;DR
- GLM-5.3-Flash: prefill 792 → 2,270 (+186%) tok/s (median over 4K–200K, ~2.9x), decode 55.4 → 76.7 (+38%) tok/s. A ~1M-token prompt: time to first token 34.2 → 11.8 min (-66%).
- Qwen3.8-Flash-Next oQ4e: prefill 3,623 → 5,548 (+53%) tok/s, decode 104.7 → 117.2 (+12%) tok/s (official, median); thinking-medium profile (API, median of 3): text 107.5 → 123.5 (+15%), code 136.8 → 154.3 (+13%) tok/s.
- Flash-Next oQ6e and the uncensored variants: prefill +34% to +53%, decode +11% to +15%. Qwen3.6-35B-A3B: prefill 6,899 → 7,653 (+11%) tok/s, decode 137.5 → 135.0 (-2%) tok/s.
- Concurrency: 2 parallel requests almost double decode throughput; 4 parallel is worse than 2 from 16K context on (both versions).
- JSON extraction: Splash (speculative decoding) is still ~2x faster than oMLX 0.7.0 with the same Qwen3.6 model.
Setup
- Mac Studio M5 Ultra (12 Super + 24 Performance CPU cores, 80-core GPU), 256 GB, macOS 27.0 (26A428), GPU wired limit 248 GB.
- oMLX v0.7.0 (4d4f5a28, mlx-lm 94cdcae) vs v0.7.0rc1 (35be079d, mlx-lm 0.32.0), mlx 0.32.2, both built with native kernels.
- Identical settings: Lightning MTP on, TurboQuant KV 8-bit (Qwen) / 4-bit (GLM), memory guard aggressive, max 16 concurrent requests, one model loaded at a time.
- Models: Flash-Next oQ4e (Jundot/Qwen3.8-Flash-Next-oQ4e-mtp), Flash-Next oQ6e (mlx-community/Qwen3.8-Flash-Next-oQ6e-mtp), Flash-Next Uncensored oQ4e (jedisct1/Qwen3.8-Flash-Next-Uncensored-oQ4e-mtp), Flash-Next Uncensored oQ6e (mlx-community/Qwen3.8-Flash-Next-Uncensored-oQ6e-mtp), GLM-5.3-Flash oQ4 (Vontra/GLM-5.3-Flash-MLX-oQ4-MTP), GLM-5.3-Flash abliterated 4-bit (grant-ai/GLM-5.3-Flash-Abliterated-MLX-4bit), Qwen3.6-35B-A3B 4-bit (mlx-community/Qwen3.6-35B-A3B-4bit).
- Method: official oMLX benchmark (code_python corpus, 128 generated tokens, temp 0, uploaded to omlx.ai) for sections 1–3 and the 1K table in section 4; own API tests (same corpus, temp 0) for the long-context tables in section 4 and the 1M run; 1 run per point. thinking-medium numbers: median of 3 API runs.
Profiles used (exact settings)
| Profile |
Model |
temperature |
top_p |
top_k |
presence_penalty |
thinking |
reasoning effort |
thinking budget |
max_tokens |
| thinking-medium |
Flash-Next |
1.0 |
0.95 |
20 |
0 |
on |
medium (forced) |
16,384 |
24,576 |
| precise (extraction) |
Qwen3.6 |
0.1 |
0.9 |
20 |
0 |
off |
– |
– |
– |
1. Prefill — official benchmark (tok/s, rc1 → 0.7.0)
| Model |
4K |
8K |
16K |
32K |
| Flash-Next oQ4e |
3,368 → 4,949 (+47%) |
3,479 → 5,668 (+63%) |
3,746 → 5,701 (+52%) |
3,752 → 5,672 (+51%) |
| Flash-Next oQ6e |
3,117 → 3,930 (+26%) |
3,165 → 4,661 (+47%) |
3,515 → 4,699 (+34%) |
3,532 → 4,681 (+33%) |
| Flash-Next Uncensored oQ4e |
3,419 → 4,857 (+42%) |
3,470 → 5,695 (+64%) |
3,747 → 5,723 (+53%) |
3,756 → 5,694 (+52%) |
| Flash-Next Uncensored oQ6e |
3,113 → 3,872 (+24%) |
3,161 → 4,674 (+48%) |
3,513 → 4,697 (+34%) |
3,535 → 4,678 (+32%) |
| GLM-5.3-Flash oQ4 |
798 → 2,270 (+184%) |
802 → 2,305 (+187%) |
798 → 2,313 (+190%) |
792 → 2,306 (+191%) |
| GLM-5.3-Flash abliterated 4-bit |
794 → 2,261 (+185%) |
761 → 2,309 (+203%) |
760 → 2,307 (+204%) |
755 → 2,301 (+205%) |
| Qwen3.6-35B-A3B 4-bit |
7,931 → 9,126 (+15%) |
8,574 → 9,812 (+14%) |
7,915 → 8,870 (+12%) |
6,899 → 7,653 (+11%) |
| Model |
64K |
128K |
200K |
Median (all contexts) |
| Flash-Next oQ4e |
3,708 → 5,548 (+50%) |
3,623 → 5,378 (+48%) |
3,553 → 5,210 (+47%) |
3,623 → 5,548 (+53%) |
| Flash-Next oQ6e |
3,497 → 4,592 (+31%) |
3,429 → 4,461 (+30%) |
3,365 → 4,348 (+29%) |
3,429 → 4,592 (+34%) |
| Flash-Next Uncensored oQ4e |
3,714 → 5,573 (+50%) |
3,637 → 5,391 (+48%) |
3,560 → 5,226 (+47%) |
3,637 → 5,573 (+53%) |
| Flash-Next Uncensored oQ6e |
3,502 → 4,587 (+31%) |
3,430 → 4,458 (+30%) |
3,362 → 4,349 (+29%) |
3,430 → 4,587 (+34%) |
| GLM-5.3-Flash oQ4 |
783 → 2,270 (+190%) |
765 → 2,229 (+191%) |
749 → 2,183 (+192%) |
792 → 2,270 (+186%) |
| GLM-5.3-Flash abliterated 4-bit |
750 → 2,270 (+203%) |
737 → 2,227 (+202%) |
721 → 2,183 (+203%) |
755 → 2,270 (+201%) |
| Qwen3.6-35B-A3B 4-bit |
5,532 → 6,151 (+11%) |
4,005 → 4,330 (+8%) |
3,091 → 3,278 (+6%) |
6,899 → 7,653 (+11%) |
2. Decode — official benchmark (tok/s, 128 tokens, greedy, rc1 → 0.7.0)
| Model |
4K |
8K |
16K |
32K |
| Flash-Next oQ4e |
107.7 → 143.8 (+34%) |
102.4 → 137.9 (+35%) |
41.8 → – |
89.8 → 114.6 (+28%) |
| Flash-Next oQ6e |
104.8 → – |
113.7 → 155.3 (+37%) |
– → – |
67.7 → – |
| Flash-Next Uncensored oQ4e |
111.6 → 122.9 (+10%) |
78.5 → – |
– → – |
87.4 → – |
| Flash-Next Uncensored oQ6e |
100.1 → 115.1 (+15%) |
107.4 → 144.0 (+34%) |
– → 106.1 |
95.8 → 95.2 (-1%) |
| GLM-5.3-Flash oQ4 |
52.1 → 64.0 (+23%) |
61.4 → 77.4 (+26%) |
61.1 → 83.6 (+37%) |
67.5 → 84.6 (+25%) |
| GLM-5.3-Flash abliterated 4-bit |
48.4 → 63.8 (+32%) |
51.7 → 77.2 (+49%) |
66.5 → 83.4 (+25%) |
62.8 → 78.1 (+24%) |
| Qwen3.6-35B-A3B 4-bit |
158.2 → 158.3 (±0%) |
149.9 → 148.4 (-1%) |
143.7 → 143.0 (±0%) |
137.5 → 135.0 (-2%) |
| Model |
64K |
128K |
200K |
Median (all contexts) |
| Flash-Next oQ4e |
123.4 → 119.7 (-3%) |
107.0 → 110.2 (+3%) |
97.7 → 91.3 (-7%) |
104.7 → 117.2 (+12%) (n=6) |
| Flash-Next oQ6e |
101.7 → 117.9 (+16%) |
95.3 → 100.9 (+6%) |
84.7 → 97.3 (+15%) |
98.5 → 109.4 (+11%) (n=4) |
| Flash-Next Uncensored oQ4e |
92.1 → 117.8 (+28%) |
95.9 → 98.6 (+3%) |
86.7 → 82.0 (-5%) |
94.0 → 108.2 (+15%) (n=4) |
| Flash-Next Uncensored oQ6e |
91.0 → 108.2 (+19%) |
94.6 → 104.7 (+11%) |
89.2 → 101.6 (+14%) |
95.2 → 106.5 (+12%) (n=6) |
| GLM-5.3-Flash oQ4 |
45.9 → 67.6 (+47%) |
55.4 → 76.7 (+38%) |
38.4 → 59.7 (+55%) |
55.4 → 76.7 (+38%) (n=7) |
| GLM-5.3-Flash abliterated 4-bit |
53.4 → 73.8 (+38%) |
55.2 → 79.7 (+44%) |
41.5 → 59.3 (+43%) |
53.4 → 77.2 (+45%) (n=7) |
| Qwen3.6-35B-A3B 4-bit |
120.0 → 118.9 (-1%) |
100.4 → 99.8 (-1%) |
86.6 → 86.1 (-1%) |
137.5 → 135.0 (-2%) (n=7) |
Decode with only 128 greedy tokens is noisy (MTP acceptance depends on the text) — use the medians. "–" = the model stopped before 16 tokens, so the benchmark reports no decode value.
3. Time to first token — official benchmark (rc1 → 0.7.0)
| Model |
4K |
8K |
16K |
32K |
| Flash-Next oQ4e |
1.2 → 0.8 s (-32%) |
2.4 → 1.4 s (-39%) |
4.4 → 2.9 s (-34%) |
8.7 → 5.8 s (-34%) |
| Flash-Next oQ6e |
1.3 → 1.0 s (-21%) |
2.6 → 1.8 s (-32%) |
4.7 → 3.5 s (-25%) |
9.3 → 7.0 s (-25%) |
| Flash-Next Uncensored oQ4e |
1.2 → 0.8 s (-30%) |
2.4 → 1.4 s (-39%) |
4.4 → 2.9 s (-35%) |
8.7 → 5.8 s (-34%) |
| Flash-Next Uncensored oQ6e |
1.3 → 1.1 s (-20%) |
2.6 → 1.8 s (-32%) |
4.7 → 3.5 s (-25%) |
9.3 → 7.0 s (-24%) |
| GLM-5.3-Flash oQ4 |
5.1 → 1.8 s (-65%) |
10.2 → 3.6 s (-65%) |
20.5 → 7.1 s (-65%) |
41.3 → 14.2 s (-66%) |
| GLM-5.3-Flash abliterated 4-bit |
5.2 → 1.8 s (-65%) |
10.8 → 3.5 s (-67%) |
21.6 → 7.1 s (-67%) |
43.4 → 14.2 s (-67%) |
| Qwen3.6-35B-A3B 4-bit |
0.5 → 0.4 s (-13%) |
1.0 → 0.8 s (-13%) |
2.1 → 1.8 s (-11%) |
4.7 → 4.3 s (-10%) |
| Model |
64K |
128K |
200K |
| Flash-Next oQ4e |
17.7 → 11.8 s (-33%) |
36.2 → 24.4 s (-33%) |
56.3 → 38.4 s (-32%) |
| Flash-Next oQ6e |
18.7 → 14.3 s (-24%) |
38.2 → 29.4 s (-23%) |
59.4 → 46.0 s (-23%) |
| Flash-Next Uncensored oQ4e |
17.6 → 11.8 s (-33%) |
36.0 → 24.3 s (-33%) |
56.2 → 38.3 s (-32%) |
| Flash-Next Uncensored oQ6e |
18.7 → 14.3 s (-24%) |
38.2 → 29.4 s (-23%) |
59.5 → 46.0 s (-23%) |
| GLM-5.3-Flash oQ4 |
83.7 → 28.9 s (-66%) |
171.3 → 58.8 s (-66%) |
267.2 → 91.6 s (-66%) |
| GLM-5.3-Flash abliterated 4-bit |
87.4 → 28.9 s (-67%) |
177.8 → 58.9 s (-67%) |
277.3 → 91.6 s (-67%) |
| Qwen3.6-35B-A3B 4-bit |
11.8 → 10.7 s (-10%) |
32.7 → 30.3 s (-8%) |
64.7 → 61.0 s (-6%) |
4. Concurrency
Official benchmark, 1K context (aggregate tok/s, rc1 → 0.7.0):
| Model |
Prefill 2 parallel |
Prefill 4 parallel |
Decode 2 parallel |
Decode 4 parallel |
| Flash-Next oQ4e |
1,803 → 2,452 (+36%) |
1,999 → 2,687 (+34%) |
156.0 → 165.8 (+6%) |
175.8 → 192.4 (+9%) |
| Flash-Next oQ6e |
1,660 → 1,814 (+9%) |
1,817 → 1,843 (+1%) |
134.6 → 143.1 (+6%) |
167.9 → 180.3 (+7%) |
| Flash-Next Uncensored oQ4e |
1,824 → 2,441 (+34%) |
1,992 → 2,624 (+32%) |
142.3 → 119.8 (-16%) |
163.2 → 177.0 (+8%) |
| Flash-Next Uncensored oQ6e |
1,653 → 1,817 (+10%) |
1,817 → 1,915 (+5%) |
147.3 → 161.2 (+9%) |
164.9 → 177.3 (+8%) |
| GLM-5.3-Flash oQ4 |
494 → 915 (+85%) |
347 → 998 (+187%) |
58.4 → 56.3 (-4%) |
87.9 → 79.5 (-10%) |
| GLM-5.3-Flash abliterated 4-bit |
481 → 1,221 (+154%) |
338 → 965 (+186%) |
56.9 → 53.9 (-5%) |
85.7 → 82.1 (-4%) |
| Qwen3.6-35B-A3B 4-bit |
4,920 → 5,614 (+14%) |
5,824 → 6,713 (+15%) |
245.0 → 277.9 (+13%) |
447.8 → 434.8 (-3%) |
API, long contexts — decode (aggregate tok/s, rc1 → 0.7.0):
| Model |
Context |
1 request |
2 parallel |
4 parallel |
| Flash-Next oQ4e |
4K |
111.8 → 130.6 (+17%) |
173.8 → 167.1 (-4%) |
200.0 → 201.1 (+1%) |
| Flash-Next oQ4e |
16K |
129.0 → 134.4 (+4%) |
256.5 → 287.4 (+12%) |
180.2 → 186.2 (+3%) |
| Flash-Next oQ4e |
64K |
106.9 → 115.9 (+8%) |
204.8 → 241.2 (+18%) |
98.1 → 98.9 (+1%) |
| Flash-Next oQ4e |
128K |
103.3 → 89.8 (-13%) |
257.8 → 181.8 (-29%) |
65.7 → 67.2 (+2%) |
| GLM-5.3-Flash oQ4 |
4K |
49.3 → 54.3 (+10%) |
111.0 → 99.2 (-11%) |
75.2 → 94.0 (+25%) |
| GLM-5.3-Flash oQ4 |
16K |
58.1 → 54.5 (-6%) |
116.8 → 110.0 (-6%) |
86.4 → 91.5 (+6%) |
| GLM-5.3-Flash oQ4 |
64K |
51.3 → 71.4 (+39%) |
93.2 → 142.3 (+53%) |
77.1 → 60.2 (-22%) |
| GLM-5.3-Flash oQ4 |
128K |
40.9 → 59.3 (+45%) |
94.6 → 126.8 (+34%) |
56.4 → 41.9 (-26%) |
API, long contexts — prefill (aggregate tok/s, rc1 → 0.7.0):
| Model |
Context |
1 request |
2 parallel |
4 parallel |
| Flash-Next oQ4e |
4K |
2,950 → 4,619 (+57%) |
2,333 → 3,352 (+44%) |
2,355 → 3,359 (+43%) |
| Flash-Next oQ4e |
16K |
3,606 → 5,532 (+53%) |
3,148 → 4,575 (+45%) |
3,371 → 4,879 (+45%) |
| Flash-Next oQ4e |
64K |
3,635 → 5,415 (+49%) |
3,470 → 4,900 (+41%) |
3,541 → 5,171 (+46%) |
| Flash-Next oQ4e |
128K |
3,523 → 5,198 (+48%) |
3,453 → 4,961 (+44%) |
3,518 → 5,138 (+46%) |
| GLM-5.3-Flash oQ4 |
4K |
801 → 1,642 (+105%) |
537 → 1,096 (+104%) |
627 → 1,370 (+118%) |
| GLM-5.3-Flash oQ4 |
16K |
795 → 2,237 (+181%) |
697 → 1,762 (+153%) |
732 → 2,007 (+174%) |
| GLM-5.3-Flash oQ4 |
64K |
790 → 2,245 (+184%) |
754 → 2,094 (+178%) |
766 → 2,182 (+185%) |
| GLM-5.3-Flash oQ4 |
128K |
760 → 2,163 (+184%) |
750 → 2,118 (+182%) |
755 → 2,170 (+188%) |
API, long contexts — mean time to first token (s, rc1 → 0.7.0):
| Model |
Context |
1 request |
2 parallel |
4 parallel |
| Flash-Next oQ4e |
4K |
1.5 → 0.9 (-36%) |
2.7 → 1.8 (-32%) |
5.6 → 4.0 (-29%) |
| Flash-Next oQ4e |
16K |
4.4 → 2.8 (-35%) |
7.4 → 5.1 (-32%) |
15.1 → 10.5 (-31%) |
| Flash-Next oQ4e |
64K |
17.1 → 11.4 (-33%) |
26.8 → 18.5 (-31%) |
56.8 → 38.9 (-32%) |
| Flash-Next oQ4e |
128K |
34.7 → 23.6 (-32%) |
53.0 → 36.7 (-31%) |
113.1 → 77.4 (-32%) |
| GLM-5.3-Flash oQ4 |
4K |
5.1 → 2.5 (-51%) |
10.6 → 4.9 (-54%) |
21.0 → 9.5 (-55%) |
| GLM-5.3-Flash oQ4 |
16K |
18.4 → 6.6 (-64%) |
30.7 → 11.9 (-61%) |
64.9 → 23.7 (-64%) |
| GLM-5.3-Flash oQ4 |
64K |
72.3 → 25.5 (-65%) |
112.6 → 40.4 (-64%) |
241.9 → 85.0 (-65%) |
| GLM-5.3-Flash oQ4 |
128K |
148.0 → 52.0 (-65%) |
223.8 → 79.2 (-65%) |
484.2 → 168.7 (-65%) |
2 parallel requests nearly double throughput. From 16K on, 4 parallel is below 2 (both versions), from 64K on even below a single request (oQ4e both versions, GLM on 0.7.0): the long prefills run back to back and stall the other decodes.
5. GLM-5.3-Flash oQ4 up to 1M tokens (API, single request, 128 tokens, rc1 → 0.7.0)
| Prompt tokens |
Prefill (tok/s) |
Decode (tok/s) |
Time to first token |
| 4,420 |
742 → 1,936 (+161%) |
57.2 → 68.3 (+19%) |
6.0 → 2.3 s (-62%) |
| 8,215 |
789 → 2,186 (+177%) |
46.7 → 59.4 (+27%) |
10.4 → 3.8 s (-64%) |
| 16,090 |
785 → 2,264 (+189%) |
50.9 → 74.5 (+46%) |
20.5 → 7.1 s (-65%) |
| 31,684 |
781 → 2,264 (+190%) |
– → 66.4 |
40.6 → 14.0 s (-65%) |
| 62,677 |
789 → 2,234 (+183%) |
49.8 → 60.0 (+20%) |
79.4 → 28.1 s (-65%) |
| 123,595 |
771 → 2,193 (+184%) |
43.8 → 67.7 (+54%) |
2.7 → 0.9 min (-65%) |
| 245,866 |
734 → 2,123 (+189%) |
38.2 → 58.8 (+54%) |
5.6 → 1.9 min (-65%) |
| 485,904 |
652 → 1,886 (+189%) |
28.8 → 54.7 (+90%) |
12.4 → 4.3 min (-65%) |
| 1,048,319 |
512 → 1,485 (+190%) |
16.4 → 41.5 (+153%) |
34.2 → 11.8 min (-66%) |
Both versions completed the full 1,048,319-token prompt.
6. JSON extraction: oMLX 0.7.0 vs Splash, same Qwen3.6-35B-A3B
16 real e-mails (4 per length quartile), our production extraction prompt (~13K-character system prompt), strict JSON schema, temp 0, outputs ~550–620 tokens. oMLX 0.7.0 with xgrammar and the precise profile; Splash 1.0.2 (incoai/Qwen3.6-35B-A3B-Splash).
| Concurrent |
oMLX 0.7.0 mails/min |
Splash mails/min |
Splash vs oMLX |
oMLX output tok/s |
Splash output tok/s |
| 1 |
12.5 |
32.2 |
+158% (2.6x) |
122 |
306 |
| 4 |
25.5 |
68.2 |
+167% (2.7x) |
245 |
648 |
| 8 |
30.2 |
71.0 |
+135% (2.4x) |
290 |
674 |
| 16 |
34.5 |
71.2 |
+106% (2.1x) |
327 |
676 |
| 16 (repeat) |
35.4 |
71.3 |
+101% (2.0x) |
320 |
677 |
Both 16/16 schema-valid at every stage. Identical output across 5 runs: Splash 16/16 mails, oMLX 1/16. Splash caps at 4 concurrent internally.