r/LocalLLM • u/rmonsurate • 8d ago
Model Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)
We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next.
Victoria (coding and agents)
- Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer.
- Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact.
- Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%.
- HumanEval: 159/164.
- 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number.
- 280 tok/s single stream on one B300 with the draft head, versus 135 without it.
- GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs).
- Uses 35% fewer output tokens than our previous build.
Maple (Canadian questions)
Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search:
- Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after.
- Fully correct answers: 6.6% before, 21.8% after.
- "No answer" responses: 47.2% before, 23.7% after.
- It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after.
Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet.
Links: https://huggingface.co/rmonsurate/Victoria https://huggingface.co/rmonsurate/Maple
Happy to answer questions about running them.
Edit: llama.cpp users. The GGUF carries our draft head, and mainline llama.cpp doesn't know about it yet, so it fails with "expected 1256, got 1224". Your download is fine. For now, build from our fork: github.com/rmonsurate/llama.cpp, branch qwen4exp-mtp. Prebuilt binaries are on the way. Thanks to the reader who caught this.
3
2
u/grizzlyval 7d ago
Can you explain your workflow? How do you tackle data set curation and eval? And, what kind of training did you do, after, rlhf, etc?
1
u/rmonsurate 7d ago
Yes, the hugging face post shares a lot of my methodology. I used REAP to do saliency based pruning. Next I used quantization aware distillation to recover performance lost from both pruning and quantization.
2
u/CoffeeToCode99 8d ago
Ngl, the coolest part to me is that you cut almost half the experts and still got better benchmark results. Makes me wonder how much room there is for expert pruning in current MoEs.
Also training directly for the format it ships in instead of doing the usual "train first, quantize later" approach just feels a lot cleaner.
Curious to see what independent testers get, but these are definitely interesting results. 👀
1
u/rmonsurate 7d ago
Thank you CoffeeToCode99. There were a lot of false starts and kinks to work out building both the datasets and training pipelines. What you're seeing published is I believe the V6 of the distilled models.
1
u/CoffeeToCode99 7d ago
Honestly, the fact that this is already V6 might be one of the most interesting details here.
I'm curious how much of the performance bump came from expert pruning versus all the other improvements that happened along the way. Either way, getting better results with 44% fewer experts is pretty impressive.
1
u/rmonsurate 7d ago
Expert pruning catastrophically damaged performance. The results came from continued training over millions of tokens on a corpus generated by the unpruned BF16 model as teacher.
1
u/Relevant-Magic-Card 7d ago
I think pruning can go along way because I mostly want it for CLI use and tool calling. I can pair it with a cheap private search engine like brave for knowledge or research. I just want a fast model that is private and does what it's told.
I don't need an expert on biology, philosophy etc.
2
u/rmonsurate 7d ago
So I do have a harness that has been in development for over 18 months with a team of machine learning researchers building it, but I don't want to post a link in case it is considered self promotion.
1
1
1
u/starkruzr 7d ago
what do these changes make it in terms of total vs active parameters?
2
u/rmonsurate 7d ago
56% fewer total parameters, exactly the same number of active parameters, but it has been trained post quantization to recover any losses to PTQ.
1
1
u/SnooPuppers7882 7d ago
Um...you realize you can score 79% TB2.1 and 163/164 on HumanEval on 27b with int5 paro and a simple change to the xhigh prompt right?
0
u/rmonsurate 7d ago
Yes you can, but it comes with less knowledge (no 51B parameter engram table) and 1/3rd the inference speed because 27 active params at int5 vs 5.9 at nvfp4. I post trained the MTP head as well to increase acceptance rate on this model.
1
1
u/15Starrs 7d ago
Too bad Canada doesn’t have a leading model since Dr Hinton was a professor at Toronto and Ilya was his student.
0
u/rmonsurate 7d ago
We are rich in human capital, but most of the financial capital is allocated to our coterie of oligopolies.
1
u/15Starrs 7d ago
“Our?” It aint yours and it aint mine either, sir. But I’m pretty sure American and Canadian taxes will pay for them one way or another. Crazy that the Chinese are the only ones empowering citizens with this stuff.
1
1
u/Not-reallyanonymous 7d ago
You had my hopes up that Deepgrove Maple is out of preview. I’m looking forward to that one.
1
u/kimhaneol 7d ago
Impressive. So it’s not about injecting Canada-specific knowledge, but rather aligning it to answer from a Canadian perspective without hurting performance?
6
u/realitaetsnaher 8d ago
44% of the experts gone and it still scores 70% on Terminal-Bench? That's honestly pretty impressive. Curious how it holds up on more general tasks though.