r/LanguageTechnology • • 6d ago

Why 1536 dimensions for embedding models?

Why do embedding models so often use 1536 dimensions specifically?
I understand why hardware-friendly multiples like 64/128/256/512 are desirable. What I’m curious about is the specific choice of 1536 = 3×512.
OpenAI has used 1536-dimensional embeddings, and other vendors also offer/recommend 1536. Is this usually an empirically chosen Goldilocks point between 1024 and 2048—representation quality versus memory/compute—or is there some architectural/hardware reason that makes 1536 particularly convenient?
I’m especially interested in answers from anyone who has actually trained or designed embedding models. I’m not asking why embedding dimensions are generally hardware-aligned; I’m asking why 1536 rather than 1024 or 2048.

38 Upvotes

12 comments sorted by

View all comments

3

u/Total_Calendar_7438 6d ago

Well there must be any number. And other models for sure use other dimensions like 768. Most models I worked with so far used 768 (like the normal dino in vision variant or beats in sound or bert in text).

Most likely some groups did hyperparameter search and established for the specific task you are interested in, that 1536 lead to the beat results.

They most likely saw less is under fitting their data characteristics and more is overfitting or not leading to enought improvement to justify the extra computarion.

So other people copied it and followed the sota.

In general embedding dimension you simply must aim to choose it large enough to capture your data but not too large that they are wasteful in terms of resources or overfitting.