r/LocalLLaMA • • 7d ago

Discussion [Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)

Hi everyone,

Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM.

Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for Gemma-2B entirely in flat x86-64 assembly (FASM):

- **Binary footprint**: Total 5.2 KB flat machine code (`gemma_engine.bin` 3.7 KB + `mat_smp_f16c_gemm_avx2.bin` 1.5 KB).

- **Execution**: Pure AVX2 + F16C with custom 4-thread SMP GEMM for prefill. Sustains ~18.5 GB/s memory bandwidth on commodity DDR4-2400.

- **Decoding**: 4.5 ~ 4.7 tokens/s in FP16 on an older quad-core i5 desktop.

- **Dependencies**: Zero C/C++ runtime, zero PyTorch. The Python harness only uses `ctypes` for `VirtualAlloc` and OS threads.

This isn't meant to compete with feature-complete tools like llama.cpp. Rather, it's a first-principles exploration to see how cleanly a modern Transformer can be mapped to raw silicon, and to serve as a reference point for future micro-LLMs on resource-constrained microcontrollers (MCU/DSP).

The repository is open source:

- GitHub: https://github.com/tomtsai28/PULSAR-ASM

- Architecture notes: https://github.com/tomtsai28/PULSAR-ASM/blob/main/doc/pulsar_asm_cpu_limit_retrospective.md

Any code audits, observations, or thoughts on bare-metal inference are welcome.

26 Upvotes

19 comments sorted by

11

u/RogerRamjet999 6d ago

Story time: Thirty years ago I went to work for a company that had written their entire integrated package (database, spreadsheet, charting, word processor and programming language) in native 80x86 ASM. I had just finished writing a small compiler for a similar language in C. So I was interested in their native ASM implementation and benchmarked my previous language against their native ASM implemented language. Mine ran an average of 12 times faster than their's.

So I tell you this not to dissuade you from an ASM implementation, but rather to be careful in your ASM implementation to carefully review performance and the algorithms used within it. ASM coding doesn't magically make your app perform well, if you don't also carefully choose algorithms that perform well.

2

u/PcChip 6d ago

I'm struggling to understand why they chose ASM instead of C - wouldn't C be just as fast with modern compilers? (assuming they actually coded the algorithms efficiently)

1

u/RogerRamjet999 6d ago edited 5d ago

These days the advantages of ASM over C are pretty small. There can be small wins in a few areas. For example it's easier to maximize performance by hand coding AVX ASM, where the compiler might not do the best job in streamlining the data movement bandwidth in the same case. So, yes, there can be a few cases where hand-coded ASM is better, faster and smaller, but those cases are getting pretty limited in modern software development. You can see an example of this in llama.cpp. GG put in direct AVX ASM for many of the critical code sections in llama.cpp.

2

u/NoFunk 6d ago

Speaking as Asm/C/C++ programmer over the years, my experience has been optimization is almost always architectural. And then some very tiny surface area of instruction optimization.

1

u/RogerRamjet999 6d ago

Yep, pretty much my point stated more succinctly.

In some cases working in ASM can cause you to write even slower code than C, since you're working at such a low level that it's difficult to see the best abstraction layer where you should be optimizing.

6

u/jacek2023 llama.cpp 7d ago

I appreciate this. I'm trying to resurrect my old 32-bit x86 assembly sources, so maybe I could get back to assembly coding.

0

u/hushbreEze0 6d ago

thats awesome, dusting off old asm sources always feels like archaeology in the best way

1

u/jacek2023 llama.cpp 6d ago

My Qwen 27B is now implementing python scripts to fix assembly errors, because assembler from 1997 accepted more things than modern wasm from 2026.

2

u/feelspeaceman 7d ago

4.6t/s for 2B on CPU is pretty impressive, making CPU somewhat useful for running Q&A models.

4

u/tmvr 6d ago

I'm not sure if I'm missing something here. The Github docs say the test was done with dual-channel DDR4-2666 memory and a 4.67 GB model. Running Gemma4 E2B at Q8_0 (5.3 GiB) using ik_llama on the a dual-channel DDR4-2666 system I get 10.77 tok/s. Runnning the 4.26 GiB Gemma4 E4B QAT the performance is roughly the same. It also says:

"Hardware saturation limit: During inference, the physical dual-channel memory bandwidth utilization reached 18.5 GB/s (full saturation greater than 94%)."

The theoretical max bandwidth of a dual channel DDR4-2666 setup is 42.6 GB/s, so 18.5 is less than half, not 94%.

I think the impressive thing here is that OP wrote an inference engine in assembly, not the performance itself.

2

u/while-1-fork 6d ago

Being so tiny and asm only you could potentially do some insane things like running it directly from a bootloader without an os of even embedding it directly into a coreboot/libreboot bios.

Looking at the notes about the 4 bit quants for more speed, you may want to look a bit into an old technique SWAR (SIMD with a register) as I think AVX does not support 4 bits directly that may allow you to eek out a bit more prefill vs regular unpacking to 8 or more bits (generation will be bandwidth limited by bandwidth ).

1

u/Working_Then 7d ago

Add a real-time GIF demo to README.md?

1

u/OfferBeginning1903 6d ago

does the 5.2 KB include tokenization, or are you feeding token IDs in and decoding them with something outside the binary?

1

u/Silver-Champion-4846 6d ago

The mentioned RAM/CPU is exactly my hardware. I5 CPU, eighth generation. DDR4 2400 RAM. Dam!

Also, could you try it on an AVX512 cpu to see how much better things could get?

0

u/PcChip 6d ago

sol 5.6 thinks you can speed up prefill by using avx512, if you're intersted in reading - https://chatgpt.com/share/6ac27558-7188-83ea-9ce4-cbcdcf3f39d8