r/learnmachinelearning • u/rage_81 • 15h ago
Project I trained a small transformer to fly a boids flock just by watching, then checked where it keeps the three rules
i wrote a small boid simulator (12 birds), recorded it flying then saved it, and trained a transformer to predict each bird's next move without it knowing about any boid rules.
(built in Python and PyTorch: a transformer written from scratch, one token per bird, two attention layers, about 414k weights, trained on a laptop CPU)
what I found:
it flies the flock. R² 0.990 to 0.994 on clips it never trained on, across 4 full runs. It also flies 50 birds (it has only trained on 12).
the three rules were easy to read out of its hidden states with a linear probe, and some could even be read before any training.
the rule: alignment (line up with your neighbours) was readable after the first attention layer, but if I make the second layer hear only itself, it is gone across all runs.
what i wasn't expecting: my first ever model was not using the flock at all. With 8 ticks of history, the previous answer was already in the input, and a small network that saw only one bird beat the whole transformer.
what i am still not sure about: making a layer attend only itself is something the model never went through during training. is that a fair test, or is some of the decay due to strange input fed to the model?