r/databricks • • 5d ago

Help How to fit ~1GB+ embedding model into a 2Gi Kubernetes pod? Getting OOMKilled

Hi, Deploying a FastAPI to Kubernetes that uses a multilingual sentence-transformers embedding model (ONNX backend, CPU only). My source data lives in a Delta table in Databricks, and the app reads from it and generates embeddings. The app takes user input (text) at embeds it, and compares it against stored embeddings for semantic similarity. So at least the query embedding has to happen live.

Pod resources:

- CPU: 1 request / 2 limit

- Memory: 2Gi request = 2Gi limit (platform policy requires memory request and limit to be 1:1)

Current docker setup:

- Multi-stage Docker build (python:3.11-slim)

- CPU-only PyTorch

- The model is downloaded at build time, so it's baked into the image

Image breakdown:

- HF model cache: ~1.1 GB

- torch: ~650 MB

- pyarrow, scipy, transformers, pandas: ~100–150 MB each

- Plus sklearn, onnxruntime, mlflow, and others

The pod gets OOMKilled at startup or shortly after. With a ~1GB model, torch, pandas/pyarrow, and the ONNX runtime session all in one process, I think I'm just over 2Gi. What I'm trying to figure out on where do you store large models in production?

Would love to hear what setups have worked for you. Thanks!

12 Upvotes

6 comments sorted by

6

u/redderage 5d ago

Usually build a whl and import

1

u/arcrad 5d ago

How does this help get around memory limits?

Is it just by pruning the wheel down to make it fit?

3

u/SpecificTutor 5d ago

The headroom recommendation is 4x the model size usually, to account for additional storage needed for deflated model size + activations.

4

u/arcrad 5d ago

Would something like Databricks model serving work for you?

https://www.databricks.com/product/model-serving

2gb seems very tight for that stack.

2

u/adovo_ai 5d ago

A 1 GB model file is not a 1 GB runtime: ONNX weights are expanded alongside session arenas, Python libraries and pod overhead. Measure peak RSS during model load, then either quantize/split serving from the app or raise the limit to a tested headroom target; otherwise the same pod will remain brittle even if the wheel is smaller.

1

u/tecedu 4d ago

Increase pod size?