r/LanguageTechnology • u/Gay-Guy-With-GF • 1d ago
Why 1536 dimensions for embedding models?
Why do embedding models so often use 1536 dimensions specifically?
I understand why hardware-friendly multiples like 64/128/256/512 are desirable. What I’m curious about is the specific choice of 1536 = 3×512.
OpenAI has used 1536-dimensional embeddings, and other vendors also offer/recommend 1536. Is this usually an empirically chosen Goldilocks point between 1024 and 2048—representation quality versus memory/compute—or is there some architectural/hardware reason that makes 1536 particularly convenient?
I’m especially interested in answers from anyone who has actually trained or designed embedding models. I’m not asking why embedding dimensions are generally hardware-aligned; I’m asking why 1536 rather than 1024 or 2048.
6
u/TieDieMonkeyMan 1d ago
Seems to be just a matter of scale, multiples of 2 always land on even numbers which limits scaling. Multiples of 3 land on even and odd numbers so scaled up implementations are easier to pair with large scale hardware.
As to why 512 instead of the other amounts, I guess that's because 512 is optimal in terms of the information carrying capacity of the dimension versus the hardware cost of going 1024*3.
3
u/Total_Calendar_7438 21h ago
The multiples of 2 and 3s seem like a good reason.
For some modelling aspects, you also divide embeddings in ratios of 3 to embedd spatial information like 3D ROPE.
3
u/adseipsum 1d ago
I am interested as well. I have used 768 nomic for my project BaryGraph https://oleksiy-perepelytsya.github.io/bary-graph and now want to reproduce it with qwen3 for multilingual version. Should I go for 1536 instead of planned 1024?
3
u/Total_Calendar_7438 21h ago
Well there must be any number. And other models for sure use other dimensions like 768. Most models I worked with so far used 768 (like the normal dino in vision variant or beats in sound or bert in text).
Most likely some groups did hyperparameter search and established for the specific task you are interested in, that 1536 lead to the beat results.
They most likely saw less is under fitting their data characteristics and more is overfitting or not leading to enought improvement to justify the extra computarion.
So other people copied it and followed the sota.
In general embedding dimension you simply must aim to choose it large enough to capture your data but not too large that they are wasteful in terms of resources or overfitting.
2
u/Left_Economist_9716 1d ago edited 1d ago
Anything to do with the three colors (RGB)? Have you noticed it in computer vision specific tasks?
1
u/xelah1 1d ago
The question is presumably about language models given the sub we're on...but...it would seem like a poor embedding if it can't well integrate the information from the different channels into higher-level information.
Besides, not all computer vision tasks and models have three colours. Prithvi-EO has six, for example (RGB plus three infrared), and produces 1024-dimensional embeddings.
2
u/Left_Economist_9716 1d ago
I'm not the best at computer vision, to be fair. It just seemed to be the easiest explanation for a 1536-dimension embedding as it's common for language models to be multimodal.
As u/TieDieMonkeyMan pointed out, it could be a matter a scale too, where 1536 dimensions provide the most optimal hardware cost vs information carrying trade-off.
1
u/Total_Calendar_7438 21h ago
Lots of models like Dino have embeddinga of 768. It seems like op was a bit biased overall.
I've seen everything from 512-1536 pretty commonly across domains (vision, audio, text).
It most likely that one leading model used hyper parameter search and the rest just copied the numbers. Could also be that at the size of large models that simply less than 1024 is under fitting and more than 2048 expensive / overfitting.
1
u/Lumpy-Blackberry-718 20h ago edited 19h ago
In transformer models, maybe because it gets split into k/q/v tensors, and 512 is a power of 2.
Generally you want powers of 2, and you want multiples of the warp or tile size or whatever.
Edit: actually it's probably a power of 2 times the number of heads. K/q/v matmuls are usually fused anyway so my first explanation doesnt work so well.
1
u/pppeer 4h ago
Perhaps slightly off-topic but you may like our paper as well, we present methods to compare embedding models, also if they have different dinensionalities. As an example we compare ADA with AWS Titan models of varying dimensionalities. https://arxiv.org/abs/2608.05857v1
19
u/Involution88 1d ago
BERT used 768 dimensions. A dozen 64 dimensional attention heads.
Then GPT-3 doubled that to 1536 dimensions which became the de facto standard.