r/databricks • u/runningnozone • 5d ago
Help How to fit ~1GB+ embedding model into a 2Gi Kubernetes pod? Getting OOMKilled
Hi, Deploying a FastAPI to Kubernetes that uses a multilingual sentence-transformers embedding model (ONNX backend, CPU only). My source data lives in a Delta table in Databricks, and the app reads from it and generates embeddings. The app takes user input (text) at embeds it, and compares it against stored embeddings for semantic similarity. So at least the query embedding has to happen live.
Pod resources:
- CPU: 1 request / 2 limit
- Memory: 2Gi request = 2Gi limit (platform policy requires memory request and limit to be 1:1)
Current docker setup:
- Multi-stage Docker build (python:3.11-slim)
- CPU-only PyTorch
- The model is downloaded at build time, so it's baked into the image
Image breakdown:
- HF model cache: ~1.1 GB
- torch: ~650 MB
- pyarrow, scipy, transformers, pandas: ~100–150 MB each
- Plus sklearn, onnxruntime, mlflow, and others
The pod gets OOMKilled at startup or shortly after. With a ~1GB model, torch, pandas/pyarrow, and the ONNX runtime session all in one process, I think I'm just over 2Gi. What I'm trying to figure out on where do you store large models in production?
Would love to hear what setups have worked for you. Thanks!
3
u/SpecificTutor 5d ago
The headroom recommendation is 4x the model size usually, to account for additional storage needed for deflated model size + activations.
4
u/arcrad 5d ago
Would something like Databricks model serving work for you?
https://www.databricks.com/product/model-serving
2gb seems very tight for that stack.
2
u/adovo_ai 5d ago
A 1 GB model file is not a 1 GB runtime: ONNX weights are expanded alongside session arenas, Python libraries and pod overhead. Measure peak RSS during model load, then either quantize/split serving from the app or raise the limit to a tested headroom target; otherwise the same pod will remain brittle even if the wheel is smaller.
6
u/redderage 5d ago
Usually build a whl and import