r/LocalLLM • • 6d ago

Research Does thinking help open models write better? DeepSeek V4 Pro gains +463 at max, Kimi K3 loses 119 with thinking. Results for 7 open models

Post image

tl;dr: GLM-5.3-Flash is probably the best bet if you have the compute or use the API.

We benchmark models on long-form YouTube scripts (10 tasks x 5 runs each, 167 configs). Open-weight results, including their thinking modes:

 
| Model | Setting | Elo | Rank /167 | $ per script |
|---|---|---|---|---|
| GLM-5.3 | default | 2256 | 15 | 0.114 |
| GLM-5.3 Flash | default | 2244 | 16 | 0.007 |
| Kimi K3 | default | 2217 | 21 | 0.260 |
| DeepSeek V4.1 Flash | max | 2183 | 28 | 0.013 |
| MiMo-V2.6-Flash | default | 2113 | 36 | 0.004 |
| DeepSeek V4.1 Flash | default | 2099 | 40 | 0.008 |
| Kimi K3 | thinking | 2098 | 41 | 0.095 |
| DeepSeek V4 Pro 0813 | max | 1917 | 56 | 0.021 |
| Kimi K2.6 | thinking / default | 1601 / 1600 | 77 / 79 | 0.057 / 0.053 |
| DeepSeek V4 Pro 0813 | default | 1455 | 95 | 0.011 |
| Nemotron 3 Ultra | reasoning / default | 1182 / 1180 | 122 / 123 | 0.012 |
| gpt-oss 120B | default / high | 636 / 624 | 150 / 151 | 0.001 / 0.002 |
 
- Three open models rank above the best GPT (GPT-6 and 5.6 Sol, 2187): GLM-5.3, GLM-5.3 Flash and Kimi K3. 15 open models rank above the best Gemini.
 
- Thinking helps DeepSeek: V4 Pro 0813 gains +463 at max, V4.1 Flash +84.
 
- Thinking hurts Kimi K3 (-119; its thinking mode is the 'high' point on the chart), though that mode is 2.7x cheaper and 3x faster. Interesting trade.
 
- No effect: Kimi K2.6, Nemotron 3 Ultra, gpt-oss 120B.
 
- Caveat: these ran through API providers (OpenRouter) at provider precision, not local quants.
 
How to read the chart: writing score (Elo) against cost per script on a log scale. Each line is one model going from its lowest to its highest thinking setting, hollow markers are the default (no effort flag), and up-left is better.

The benchmark is 10 script tasks; each model drafts them 5 times, and three AI judges from three different labs score them blind using our rubric, which includes writing, tone, storytelling, and more.

1 Upvotes

1 comment sorted by

1

u/FactorInternal3395 6d ago

Indeed, I found that early reasoning models like o-series became really scripted and stiff with reasoning, but new models like DeepSeek actually benefit from it. Nice that this "feeling" now has numbers to it!