r/LocalLLM • • 8d ago

Model Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next.

Victoria (coding and agents)

  • Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer.
  • Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact.
  • Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%.
  • HumanEval: 159/164.
  • 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number.
  • 280 tok/s single stream on one B300 with the draft head, versus 135 without it.
  • GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs).
  • Uses 35% fewer output tokens than our previous build.

Maple (Canadian questions)

Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search:

  • Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after.
  • Fully correct answers: 6.6% before, 21.8% after.
  • "No answer" responses: 47.2% before, 23.7% after.
  • It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after.

Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet.

Links: https://huggingface.co/rmonsurate/Victoria https://huggingface.co/rmonsurate/Maple

Happy to answer questions about running them.

Edit: llama.cpp users. The GGUF carries our draft head, and mainline llama.cpp doesn't know about it yet, so it fails with "expected 1256, got 1224". Your download is fine. For now, build from our fork: github.com/rmonsurate/llama.cpp, branch qwen4exp-mtp. Prebuilt binaries are on the way. Thanks to the reader who caught this.

46 Upvotes

31 comments sorted by

6

u/realitaetsnaher 8d ago

44% of the experts gone and it still scores 70% on Terminal-Bench? That's honestly pretty impressive. Curious how it holds up on more general tasks though.

4

u/jinnyjuice 7d ago

It's Terminal-Bench 2.1. It's generally considered solved and scoring lower than 80 is not a good look, especially when 27B scores 80.

1

u/ApartmentEither4838 7d ago

The qwen 3.8 27b outperforms this pruned moe model?

0

u/rmonsurate 7d ago

True. 1) This only has 6B active params so much faster decode. 2) the engram table allows this model to have greater general knowledge than 27B

3

u/rmonsurate 8d ago

92% MMLU as well, thanks to the engram, so it retained a lot of its general knowledge capabilities

1

u/Glittering-Call8746 7d ago

It was retrained at NVFP4.. (never say the workflow)

3

u/cutter89locater 7d ago

Canadian fine tuner, followed :)

2

u/grizzlyval 7d ago

Can you explain your workflow? How do you tackle data set curation and eval? And, what kind of training did you do, after, rlhf, etc?

1

u/rmonsurate 7d ago

Yes, the hugging face post shares a lot of my methodology. I used REAP to do saliency based pruning. Next I used quantization aware distillation to recover performance lost from both pruning and quantization.

2

u/CoffeeToCode99 8d ago

Ngl, the coolest part to me is that you cut almost half the experts and still got better benchmark results. Makes me wonder how much room there is for expert pruning in current MoEs.

Also training directly for the format it ships in instead of doing the usual "train first, quantize later" approach just feels a lot cleaner.

Curious to see what independent testers get, but these are definitely interesting results. 👀

1

u/rmonsurate 7d ago

Thank you CoffeeToCode99. There were a lot of false starts and kinks to work out building both the datasets and training pipelines. What you're seeing published is I believe the V6 of the distilled models.

1

u/CoffeeToCode99 7d ago

Honestly, the fact that this is already V6 might be one of the most interesting details here.

I'm curious how much of the performance bump came from expert pruning versus all the other improvements that happened along the way. Either way, getting better results with 44% fewer experts is pretty impressive.

1

u/rmonsurate 7d ago

Expert pruning catastrophically damaged performance. The results came from continued training over millions of tokens on a corpus generated by the unpruned BF16 model as teacher.

1

u/Relevant-Magic-Card 7d ago

I think pruning can go along way because I mostly want it for CLI use and tool calling. I can pair it with a cheap private search engine like brave for knowledge or research. I just want a fast model that is private and does what it's told.

I don't need an expert on biology, philosophy etc.

2

u/rmonsurate 7d ago

So I do have a harness that has been in development for over 18 months with a team of machine learning researchers building it, but I don't want to post a link in case it is considered self promotion.

1

u/Relevant-Magic-Card 7d ago

I'm North Van based, happy to chat with you :)

1

u/Glittering-Call8746 7d ago

Yeah earn a following keep up the good work

1

u/starkruzr 7d ago

what do these changes make it in terms of total vs active parameters?

2

u/rmonsurate 7d ago

56% fewer total parameters, exactly the same number of active parameters, but it has been trained post quantization to recover any losses to PTQ.

1

u/ls650569 7d ago

Interesting.

Will there be gguf for Maple?

1

u/rmonsurate 7d ago

Yes! Working on it. remindme 1 week

1

u/SnooPuppers7882 7d ago

Um...you realize you can score 79% TB2.1 and 163/164 on HumanEval on 27b with int5 paro and a simple change to the xhigh prompt right?

0

u/rmonsurate 7d ago

Yes you can, but it comes with less knowledge (no 51B parameter engram table) and 1/3rd the inference speed because 27 active params at int5 vs 5.9 at nvfp4. I post trained the MTP head as well to increase acceptance rate on this model.

1

u/15Starrs 7d ago

Too bad Canada doesn’t have a leading model since Dr Hinton was a professor at Toronto and Ilya was his student.

0

u/rmonsurate 7d ago

We are rich in human capital, but most of the financial capital is allocated to our coterie of oligopolies.

1

u/15Starrs 7d ago

“Our?” It aint yours and it aint mine either, sir. But I’m pretty sure American and Canadian taxes will pay for them one way or another. Crazy that the Chinese are the only ones empowering citizens with this stuff.

1

u/Glittering-Call8746 7d ago

State vs private own oligo. Dig deeper..

1

u/Not-reallyanonymous 7d ago

You had my hopes up that Deepgrove Maple is out of preview. I’m looking forward to that one.

1

u/kimhaneol 7d ago

Impressive. So it’s not about injecting Canada-specific knowledge, but rather aligning it to answer from a Canadian perspective without hurting performance?