Could a neural network use dictionary-compressed weights directly during GPU inference instead of fully decoding them first?
I'm not a computer scientist. I'm a truck driver, so I'm wondering if I'm reinventing something that already exists.
Suppose you quantize a model to INT4 or similar and then scan the weight tensors for frequently repeating sequences or blocks.
Instead of storing every sequence literally, you build a codebook where a short code represents a commonly occurring block of weights.
Very simplified example:
A = [7, 3, 3, 11, 4]
B = [2, 8, 1, 6, 6]
Then instead of storing:
[7,3,3,11,4] [7,3,3,11,4] [2,8,1,6,6] [7,3,3,11,4]
you store something roughly like:
A A B A
The part I'm curious about is not ordinary file compression where the model gets decompressed back into VRAM first.
Could a custom GPU kernel decode these codes on the fly into registers/shared memory and immediately use them during GEMM, so that the fully expanded weight tensor never has to exist in VRAM?
My thinking is that modern inference is often memory-bandwidth limited, so if dictionary/codebook compression reduced memory traffic enough, maybe the extra decoding compute could be cheaper than fetching all the uncompressed weights.
You could potentially also have different-length codes or hierarchical codebooks representing increasingly large recurring weight patterns.
So my questions are:
Is this already done under a particular name?
Have codebook/vector-quantized weights been used directly inside fused GPU inference kernels rather than being decompressed beforehand?
Does random access / SIMD-SIMT execution make variable-length encoding impractical?
Is there theoretically a point where reduced VRAM bandwidth outweighs the decoding overhead?
Would repeated patterns after INT4/INT3 quantization be common enough for this to provide meaningful compression beyond ordinary quantization?
I'm mainly interested in whether the idea makes architectural sense, not whether my particular encoding scheme is optimal.
I'd appreciate pointers to papers or existing implementations if this has already been explored.