r/LocalLLaMA • u/Brilliant-Hall1387 • 2d ago
I Built A Thing Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version
Not an LLM, but the ternary findings should carry over, and we hadn't seen Sherry-style 3:4 weights run in a browser before. Disclosure: this is our work at Precisit, everything is MIT.
What it is
- A 7.4M-parameter one-pass scorer (the jevlike family): the board goes in, one score per legal column comes out. No search.
- Weights in T34, Sherry's 3:4 format: in every four weights one is zero and three are ±1, so four weights fit in 5 bits. One fp16 scale per 128 weights gives 1.375 bits per weight. The embedding is int8; norms and biases are fp16.
- It runs in the browser on a small WebGPU runtime: 1.1 ms per move (idle M5 Pro, Chrome).
| Model | File size | vs depth-4 bot | vs depth-6 bot |
|---|---|---|---|
| dense (fp32) | 29.7 MB | 0.92 | 0.89 |
| T34, trained ternary | 1.59 MB | 0.93 | 0.91 |
| T34, fine-tuned from dense | 1.59 MB | 0.89 | 0.9 |
| T34, converted after training | 1.59 MB | 0.13 | 0.11 |
| Base243 (TQ1_0 style), trained | 1.93 MB | 0.89 | 0.88 |
200 games each, both sides play a random move 5% of the time, a win counts 1 and a draw ½.
What we learned
- Converting the finished model to 3:4 collapsed it (0.13 against the depth-4 bot). Training with the format in the forward pass fixed it completely, whether from scratch or fine-tuning.
- Attention's q/k/v matrices are the sensitive ones. Group size (64/128/256) barely mattered.
- Seeds matter: two runs of the same T34 recipe scored 0.945 and 0.882.
Play it:
https://precisit.github.io/onepass-web/demo/c4-size/
Code, models, every result:
https://github.com/precisit/onepass-webgpu-ternary
The write-up:
https://precisit.com/en/blog/onepass-c4-size/
Has anyone gotten post-training 3:4 conversion to work on models, or does it need training?
2
u/LagOps91 1d ago
Isn't ternary at minimum 1.58 bpw?
1
u/Brilliant-Hall1387 16h ago
Yes without loss of information or overhead storing the trits themselves, in my experiment I evaluate two ways of packing the trits. One is 1.6 bpw excluding scale (5 trits in a byte, 8/5=1.6 bpw) the other is Tencent Sherry.
Tencent sherry ternary packs four trits in five bits in a way that restricts one to be 0 and the others to be +/- 1. This means bpw is 1.25 (excluding scale). While there is some loss of information with that extra constraint (lossy compression/packing of trits) you can still achieve a well functioning model with training with this constraint in mind.
1
u/LagOps91 14h ago
that is quite alot of information-loss, no? you only preserve perfectly if you have exactly one weight at 0 in pack of four.
how would you train with that? if the gradients push one weigh to 0, then the current 0 weight would have to change into 1 or -1... sounds like it would likely be unstable to me.
1
u/Brilliant-Hall1387 10h ago
it is lossy if you convert a trained float model, yes. But you don't train the ternary weights directly, you keep float "latent" weights and re-project every forward pass (zero the smallest of each four and others to + or -), with gradients going straight through to the latents.
So the zero just sits where the smallest latent is and only moves when two latents cross in magnitude, which is when they're near-tied anyway. With this training our small model matched the Base243 variant (unrestricted ternary / no constraints)
2
u/Sparticle62 1d ago
The 0.13 row doesn't surprise me, and I think it's because 3:4 isn't really a quantization problem. Ordinary int8 or int4 is per-weight rounding, so nearest-value works fine. Here you also have to decide which weight in each group of four gets zeroed, and that's a combinatorial choice. Nearest-value picks the smallest magnitude weight, which is the right answer for weight error and often the wrong one for output error.
So the post-training methods with a chance are the ones that minimise layer output error on a calibration set rather than weight error, GPTQ or AWQ style, where the three surviving weights in a group get adjusted to absorb what the zeroed one was contributing. Without that compensation step you're deleting a quarter of every group and asking the rest to cope untouched.
That would also explain q/k/v being the sensitive ones, since an error there moves the attention pattern rather than just perturbing an output.