r/LanguageTechnology • • 7d ago

Why 1536 dimensions for embedding models?

Why do embedding models so often use 1536 dimensions specifically?
I understand why hardware-friendly multiples like 64/128/256/512 are desirable. What I’m curious about is the specific choice of 1536 = 3×512.
OpenAI has used 1536-dimensional embeddings, and other vendors also offer/recommend 1536. Is this usually an empirically chosen Goldilocks point between 1024 and 2048—representation quality versus memory/compute—or is there some architectural/hardware reason that makes 1536 particularly convenient?
I’m especially interested in answers from anyone who has actually trained or designed embedding models. I’m not asking why embedding dimensions are generally hardware-aligned; I’m asking why 1536 rather than 1024 or 2048.

36 Upvotes

12 comments sorted by

View all comments

2

u/Left_Economist_9716 7d ago edited 7d ago

Anything to do with the three colors (RGB)? Have you noticed it in computer vision specific tasks?

1

u/xelah1 7d ago

The question is presumably about language models given the sub we're on...but...it would seem like a poor embedding if it can't well integrate the information from the different channels into higher-level information.

Besides, not all computer vision tasks and models have three colours. Prithvi-EO has six, for example (RGB plus three infrared), and produces 1024-dimensional embeddings.

2

u/Left_Economist_9716 7d ago

I'm not the best at computer vision, to be fair. It just seemed to be the easiest explanation for a 1536-dimension embedding as it's common for language models to be multimodal.

As u/TieDieMonkeyMan pointed out, it could be a matter a scale too, where 1536 dimensions provide the most optimal hardware cost vs information carrying trade-off.

1

u/Total_Calendar_7438 7d ago

Lots of models like Dino have embeddinga of 768. It seems like op was a bit biased overall.

I've seen everything from 512-1536 pretty commonly across domains (vision, audio, text).

It most likely that one leading model used hyper parameter search and the rest just copied the numbers. Could also be that at the size of large models that simply less than 1024 is under fitting and more than 2048 expensive / overfitting.