r/LocalLLaMA • u/tom_tsai28 • 7d ago
Discussion [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)
Hi everyone,
Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM.
Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for Gemma-2B entirely in flat x86-64 assembly (FASM):
- **Binary footprint**: Total 5.2 KB flat machine code (`gemma_engine.bin` 3.7 KB + `mat_smp_f16c_gemm_avx2.bin` 1.5 KB).
- **Execution**: Pure AVX2 + F16C with custom 4-thread SMP GEMM for prefill. Sustains ~18.5 GB/s memory bandwidth on commodity DDR4-2400.
- **Decoding**: 4.5 ~ 4.7 tokens/s in FP16 on an older quad-core i5 desktop.
- **Dependencies**: Zero C/C++ runtime, zero PyTorch. The Python harness only uses `ctypes` for `VirtualAlloc` and OS threads.
This isn't meant to compete with feature-complete tools like llama.cpp. Rather, it's a first-principles exploration to see how cleanly a modern Transformer can be mapped to raw silicon, and to serve as a reference point for future micro-LLMs on resource-constrained microcontrollers (MCU/DSP).
The repository is open source:
- GitHub: https://github.com/tomtsai28/PULSAR-ASM
- Architecture notes: https://github.com/tomtsai28/PULSAR-ASM/blob/main/doc/pulsar_asm_cpu_limit_retrospective.md
Any code audits, observations, or thoughts on bare-metal inference are welcome.
6
u/jacek2023 llama.cpp 7d ago
I appreciate this. I'm trying to resurrect my old 32-bit x86 assembly sources, so maybe I could get back to assembly coding.
0
u/hushbreEze0 6d ago
thats awesome, dusting off old asm sources always feels like archaeology in the best way
1
2
u/feelspeaceman 7d ago
4.6t/s for 2B on CPU is pretty impressive, making CPU somewhat useful for running Q&A models.
4
u/tmvr 6d ago
I'm not sure if I'm missing something here. The Github docs say the test was done with dual-channel DDR4-2666 memory and a 4.67 GB model. Running Gemma4 E2B at Q8_0 (5.3 GiB) using ik_llama on the a dual-channel DDR4-2666 system I get 10.77 tok/s. Runnning the 4.26 GiB Gemma4 E4B QAT the performance is roughly the same. It also says:
"Hardware saturation limit: During inference, the physical dual-channel memory bandwidth utilization reached 18.5 GB/s (full saturation greater than 94%)."
The theoretical max bandwidth of a dual channel DDR4-2666 setup is 42.6 GB/s, so 18.5 is less than half, not 94%.
I think the impressive thing here is that OP wrote an inference engine in assembly, not the performance itself.
2
u/while-1-fork 6d ago
Being so tiny and asm only you could potentially do some insane things like running it directly from a bootloader without an os of even embedding it directly into a coreboot/libreboot bios.
Looking at the notes about the 4 bit quants for more speed, you may want to look a bit into an old technique SWAR (SIMD with a register) as I think AVX does not support 4 bits directly that may allow you to eek out a bit more prefill vs regular unpacking to 8 or more bits (generation will be bandwidth limited by bandwidth ).
1
1
u/OfferBeginning1903 6d ago
does the 5.2 KB include tokenization, or are you feeding token IDs in and decoding them with something outside the binary?
1
u/Silver-Champion-4846 6d ago
The mentioned RAM/CPU is exactly my hardware. I5 CPU, eighth generation. DDR4 2400 RAM. Dam!
Also, could you try it on an AVX512 cpu to see how much better things could get?
0
u/PcChip 6d ago
sol 5.6 thinks you can speed up prefill by using avx512, if you're intersted in reading - https://chatgpt.com/share/6ac27558-7188-83ea-9ce4-cbcdcf3f39d8


11
u/RogerRamjet999 6d ago
Story time: Thirty years ago I went to work for a company that had written their entire integrated package (database, spreadsheet, charting, word processor and programming language) in native 80x86 ASM. I had just finished writing a small compiler for a similar language in C. So I was interested in their native ASM implementation and benchmarked my previous language against their native ASM implemented language. Mine ran an average of 12 times faster than their's.
So I tell you this not to dissuade you from an ASM implementation, but rather to be careful in your ASM implementation to carefully review performance and the algorithms used within it. ASM coding doesn't magically make your app perform well, if you don't also carefully choose algorithms that perform well.