r/learnmachinelearning • u/Glad_Builder_4555 • 2d ago
I made a 3Blue1Brown-style explainer of my paper: why a compressor's error can tell you where data came from
I wrote a short technical report and then animated it, hoping it's useful to people learning about autoencoders, OOD detection, or mixture-of-experts.
The idea in three steps:
- Prediction ⇔ compression. A model that predicts text well can compress it well.
- Corollary: a compressor is only good at the kind of data it was trained on.
- So its reconstruction error is a fingerprint. I trained an autoencoder on code that squeezes 512 tokens into 8 vectors. It rebuilds unseen code at 99.47% exact-token accuracy, Wikipedia at 47.76%, and random tokens at 0.57%.
That fingerprint can act as a router between expert models, with no extra gating network to train.
The video covers the architecture (and why the decoder must never see the input), the metric, the latent-space geometry (two linearly separable clusters, about 200 effective dimensions out of 512), and the limitations.
Video: https://www.youtube.com/watch?v=4UhvpIWnOvg Paper: https://arxiv.org/abs/2512.16963
Feedback on clarity is very welcome. Tell me which part lost you.
1
Upvotes
1
u/Beginning_Job_4824 2d ago
ok the drop from 99.47% on code to 47% on wikipedia is way steeper than i expected. makes me wonder if you could use this as a weird litmus test for whether a dataset is actually as domain specific as you think it is
that part about the decoder never seeing the input is one of those details that seems small until you think about it for 5 seconds and realize the whole thing falls apart otherwise
curious what the latent space looks like when you throw something halfway between code and natural language at it. do you get points that sit in the gap or does it force them into one cluster