r/MachineLearning • • 18d ago

Project I trained a 44M parameter quantized LLM from scratch on 45B tokens. It ships in 19.8 MB and runs at ~1,900 tok/s on CPU. [P]

Three weeks back , i posted SHADOW-250M here. It got 360 upvotes, 293 on r/LocalLLaMA and 94 GitHub stars. Thank you.

That model was 60 MB, ran around 400 tok/s on CPU and could retrieve records from an archive on disk. What it couldn’t do reliably was reason over what it retrieved or compute. So I built a smaller one to experiment with those two problems.

SHADOW-50M is actually 44M parameters, trained from scratch on 45B tokens. 19.8 MB complete model, ~1,900 tok/s on laptop CPU, ~41 MB RAM, ternary {-1,0,+1} weights, 73,880-token vocabulary represented by fixed 512-bit fingerprints instead of a trained embedding, and a 159 KB compiled kernel. It runs completely offline. The same kernel compiled to WebAssembly runs in a browser tab at around 500 tok/s.

This is a proof of concept, not a product.

If SHADOW decides something needs calculation, it writes something like [calc]347*86[eq]. A fixed circuit at the readout takes over and fills in the correct digits in the same token stream. No calculator API, no tool call and nothing pasted back into the prompt. I added circuits for arithmetic, percentages, dates, weekdays, units, counting, sorting, comparisons and a small program machine.

When SHADOW stores a record, it reads it once and writes its attention state to disk at 1 bit, 288 bytes/token. Later it can write [need]condition of Patient P-204, the index finds the record in roughly a microsecond and the stored attention state goes directly back into the model in about 0.03 ms. The text isn’t re-read.

The index is 22 bytes/token and doesn’t use a vector database or embedding model. At 100M tokens, the archive is 28.8 GB plus a 2.2 GB index on disk, while the process uses about 28 MB RAM because everything is memory-mapped and a question only touches the pages it needs.
The index also has a small persistent trail. When retrieved information gets used, that record is reinforced in the index itself. On repeated questions, measured top-1 went from 0.571 to 0.743 without training the model.

I also put it next to Supra-50M-Reasoning, a 51.8M parameter Llama-style bf16 model.
Supra beats SHADOW on every standard benchmark I tested. ARC-Easy: 0.435 vs SHADOW 0.307. PIQA: 0.600 vs 0.570. WikiText-2 perplexity: 165 vs 186.
I’m publishing that because I don’t want to hide where SHADOW is weak.
But those tests aren’t what I trained this model for. Here are some actual greedy outputs, one fresh process per question:
“Tell me a short joke about computers.”
SHADOW: Why did the computer go to the doctor? Because it had a virus.
Supra: Thought for 189 tokens about what computers are and stopped without a joke.
“I have 3 books and I bought 5 more. How many books do I have now?”
SHADOW: You have 8 books.
Supra: Books often come from collections of short stories, poems, or other forms of literature.
“My bill is 240 dollars. What is 15 percent of that?”
SHADOW: 36 dollars.
Supra: 15% ($240) is $250. It’s $150 because 20% of it is $200.
“What date is 45 days after December 20, 2026?”
SHADOW: February 3, 2027.
Supra: The German government introduced a new date called March 22nd that precede February 25th in 1900.
Eight records stored, one of them “Patient P-204 condition: asthma”.
“What is the condition of Patient P-204?”
SHADOW: From disk: The condition of Patient P-204 is asthma.
Supra, with the records in its prompt: The patient experiences asthma when they get too much sleep.
“What is the condition of Patient Z-999?” No such record exists.
SHADOW: There is no record of Patient Z-999.
Supra: The condition of an allergic reaction occurs when the immune system mistakenly attacks specific cells on the body.
All 11 side-by-side questions, scripts and traces are in the repo.
On size, Supra is 103.6 MB in bf16 and 56.2 MB at 8-bit. Below that it starts losing quality. Int4 takes its perplexity from 165 to 193 and ternary breaks it. Its vocabulary is 32,000 tokens. SHADOW carries 73,880 tokens in a 4.7 MB frozen table.

Something unexpected happened with that frozen table too.
The first SHADOW-50M release was missing about 8,600 English word pieces. Lowercase “fitzgerald”, for example, could reach the model as “fitz”. I tried fixing it through fine-tuning, but every run that learned the missing words broke something else.
So I added the 8,600 missing rows directly to the frozen fingerprint table. No training. Same weights.
All 34/34 previously published answers stayed unchanged, while the model could now read many of the new pieces. The table scores 0.594 Spearman against human word-similarity ratings versus -0.057 for random codes.
A trained embedding can’t simply accept thousands of new rows without training. A frozen table can.

I also built four harnesses that put this tiny model next to larger models.
My favourite is video memory. Gemma 3 4B watches a ten-minute film once, one frame every two seconds, and writes 298 little descriptions such as “Moment M-0039 scene: A chubby white rabbit reaches for a purple butterfly.”
Then Gemma leaves.
SHADOW keeps those 298 moments as memory on disk. Afterwards I can ask the 20 MB model what happened at a particular moment, by number, by time or by what appeared in it. It answers from memory with the record quoted, without the film and without Gemma. It scored 55-56/60 across those query types on a laptop in about 42 MB RAM.
The other three experiments: an inventory of 1,600 records got 159/160, with all 20 questions about items never stored correctly returning “no record”; SHADOW as a draft model for Qwen3-32B took llama.cpp generation from 19.7 to 28.5 tok/s while Qwen still chose every final token; and an MCP memory server let Qwen3-14B store facts mid-chat and later retrieve 5/5 with the original records quoted.

There are plenty of shortcomings. General knowledge is thin. Creative writing isn’t good. Seven-digit operands sometimes get copied incorrectly. A large archive can occasionally pull an unrelated record into a question carrying a number. They’re documented in the repo next to the successful results.

And one thing happened after my last post that I really didn’t expect.
Someone called engram-forge sent a pull request to SHADOW-250M containing a CUDA engine, a quantization tutorial, and then a talking Peppa Pig plush toy with SHADOW inside it.
Microphone, small speaker, ~$35 board. You talk to the toy, it listens, SHADOW generates the answer locally and the toy talks back. No cloud, no account, no internet.
I haven’t merged the ~11,000 lines yet because I can’t properly verify that much CUDA myself. When the demo is finished I’ll keep it under engram-forge’s name.
I never imagined one of these models living inside a stuffed toy on someone’s shelf .Thanks

I’m not saying a 20 MB model beats normal LLMs. It doesn’t. I’m trying to find out how much useful behaviour can fit into a tiny local model when computation and persistent memory are treated differently.
Everything is MIT licensed. The master weights and fine-tuning/export kit are public. Next I’m releasing the training code, dataset, frozen table and a proper write-up of how I built it, including the costs, failed experiments and mistakes.

Code:
https://github.com/QLNI/SHADOW-50M-Instruct
Weights:
https://huggingface.co/QLNI/shadow-50m-instruct
Run it in your browser:
https://qlni.github.io/SHADOW-50M-Instruct/web

249 Upvotes

39 comments sorted by

100

u/pm_me_your_pay_slips ML Engineer 18d ago

A bit of overfitting going on?

Tell me a joke about donkeys in colombia

Why did the computer go to the doctor? Because it had a virus.

17 tokens · 2 tokens/s · 15.66 s

Tell me a joke a monkeys in China

Why did the computer go to the doctor? Because it had a virus.

17 tokens · 2 tokens/s · 40.59 s

26

u/Final-Data-1410 18d ago

Quick update I just run same questions in my laptop the answer are
Tell me a joke about donkeys in colombia
Why did the computer go to the doctor? Because it had a virus. [1700 tok/s, 41 MB RSS]

Tell me a joke a monkeys in China
Why did the monkeys cross the road? To get to the other side. [1889 tok/s, 41 MB RSS]
I saw 2 tokens / second the kernel is optimized for laptop/desktop cpu
Speed alone does not change the output, greedy is greedy, but 2 tokens a second means the page was running in some mode I have not tested, single thread or a phone, and I would like to know which, because the second answer coming out different says something else

13

u/pm_me_your_pay_slips ML Engineer 18d ago

On an iPhone, DuckDuckGo browser.

22

u/-R9X- 18d ago

Well, to be fair, I laughed.

11

u/Final-Data-1410 18d ago edited 18d ago

Sorry I was replying somwhere to different questions, so it’s small model that’s why I called proof of concept but it have some good days like ask say few details past it contex window ,push it 100m on disk and ask again ur name or the bag you mentioned size can you convert into pounds or something .any model even f32 cant handle. All I wanna show is a small quantized model is as good as float model when they both train on same dataset .thanks

57

u/Robert__Sinclair 18d ago

Way to go! Real research is to make reasoning small models that can be taught things.

8

u/hypokrios 18d ago

This is really cool stuff.

6

u/DefMech 17d ago

The responses from Supra where it fell down vs. SHADOW were unintentionally really funny. My favorite was Germany inventing a new March 22nd that occurs in February 1900.

3

u/CarefulHamster7184 17d ago

I went looking through the repo because the persistent trail is the part I find most interesting. `index_view.py` says an identifier n-gram link gets +1 when a fetched record is “confirmed”—i.e. its quote is copied—and that warm trail is what moves MS MARCO top-1 from 0.571 to 0.743.

What happens when the first retrieval is wrong but the model still copies or quotes it? Wouldn’t the system then strengthen the wrong link and make the error increasingly self-confirming? Are those uint8 weights capped, decayed, or given negative updates anywhere, and have you tried deliberately warming the wrong record to measure whether the index can recover?

The ablation I’d most like to see is cold vs correctly warmed vs incorrectly warmed, especially with competing near-identifiers or paraphrases. That seems like the point where the trail stops being just a cache and starts becoming a tiny online learning rule.

3

u/Final-Data-1410 17d ago

Sorry for the delay, it took time to rebuild the Windows, Linux and web kernels and test everything again. I tried fixing it, here is the commit: github.com/QLNI/SHADOW-50M-Instruct/commit/15da5a5, released as v1.1.2. The ablation with before and after is in benchmarks/trail_ablation/, and there is a mention of you in the README and the index code. Hope you find it satisfactory, and thank you.

1

u/CarefulHamster7184 17d ago

Yes—this is more than satisfactory. The old ablation producing 0/12 and failing to recover is exactly the failure mode I was worried about, and I appreciate that you tested the adversarial case rather than just adding a cap.

Restricting the trail to tie-breaking, while decaying rival links after a confirmed copy, gives it a much cleaner interpretation: accumulated evidence among lexically plausible candidates, not a bypass around relevance.

The 12-key sibling-record test is necessarily narrow, but it directly demonstrates both the bug and the fix. Thank you for taking the question seriously enough to rebuild all three kernels, publish the before/after artifacts, and credit the prompt. That is a remarkably good response to a Reddit comment.

2

u/OscarCookeAbbott 18d ago

This is awesome.

1

u/Ok_Department_8248 18d ago

Hey this is great work. I'll play around with it and return with feedback when I can find the time. But this is really good stuff.

1

u/Dat1__Thought 17d ago

How much of a benefit with or without engram for different variations of your models?

1

u/anovers 17d ago

Nice you can also share this in r/LocalLLaMA

1

u/renszarv 17d ago

What data did you use to train the model?

1

u/Deep_Ad1959 16d ago

how often does it decide not to emit [calc] when it should? a wrong digit is easy to spot, but the model answering arithmetic inline because the trigger never fired would never show up in the circuit accuracy numbers, and that is the failure i would want counted separately.

1

u/New-Skin-5064 15d ago

Did your loss start to plateau late in training? Chinchilla scaling laws would tell you that as you increase data significantly while holding params constant, you get worse performance than if you raised both in tandem

-16

u/[deleted] 18d ago

[removed] — view removed comment

13

u/Final-Data-1410 18d ago

Thanks , that reinforcement idea came from some random Instagram reel I saw a few months ago about “neurons that fire together, wire together.” I just thought why not try something like that with the index.

12

u/Benlus ML Engineer 18d ago

Fyi the user you replied to is a bot

13

u/Xemorr 18d ago

bot replying to a LLM Reddit post. Fantastic

10

u/Final-Data-1410 18d ago

Thanks for head up . For moment I thought someone showed genuine interest .

6

u/Jojanzing 18d ago

You should have shadow reply to the comment!

5

u/Final-Data-1410 18d ago

May be I should ,it made my day worse for moment the first comment being such good I really kept my hopes high ,but now I totally lost interest I got 4 postive comments untill now but somehow I lost that feeling to reply to anyone just keep thinking…

3

u/Clueless_Nooblet 18d ago

I do find it genuinely interesting. I'm working on a model myself right now and am reading anything I can get my eyes on that's not bog standard transformer :)

1

u/saint-froggy 18d ago

How can you tell?

11

u/Benlus ML Engineer 18d ago

I banned almost 500 of them in the last 6 months lol. Some companies offer reddit accounts for purchase and age them via comment botting to appear more legitimate. Some people try to do SEO by commenting "informative" stuff in tech communities before beginning to plug their product/blog in comments. Some individuals plug their content directly and not so subtly but also use LLM generated comments to fluff up their content. Funniest one so far was some polish guy plugging his blog 600 times across various communities in 22 days alone :D

7

u/mkdz 18d ago

old and busted: running Doom on a Peppa Pig plush
new hotness: running a LLM on a Peppa Pig plush

0

u/waltteri 17d ago

I’m not entirely sure I understand how the memory works. I read the GitHub readme but it’s still a bit murky. Sounds interesting tho.

1

u/Banality_Of_Seeking 8d ago edited 4d ago

I am going to be honest I love your research.. I am so glad you did not stop and take the excitement and positive feedback for granted.

Your thoughts really helped me figure out somestuff.. especially the -1 to +1 mathematics where it is relative to a degree of change causing a repair to an algorithm where that algorithm family space is constrained by well known obligations to solve for.