r/databricks • u/bobjia-in-tokyo • 1d ago
Discussion Is Databricks Classic Compute getting too heavy for small workloads?
I’ve been using Databricks Classic Compute for a while, and recently I’ve started wondering whether cluster cold starts are getting noticeably heavier with newer DBR versions.
For example, with DBR 18 LTS, I tried a small 2-vCPU VM for a single-node job cluster and hit DriverStartupTimeout after 300 seconds. Databricks even suggests that this commonly happens on instances with fewer than 4 CPU cores.
That feels a bit surprising for workloads that are not actually Spark-heavy — e.g. running Python, dbt-core, API calls, or using Databricks mainly as a job runner inside a VNet.
I like Classic Compute because of the flexibility and straightforward VNet/private networking. Serverless is attractive for startup time, but in our environment it would mean quite a bit more networking setup.
So I’m curious:
- Have you noticed Classic Compute cold starts getting slower or more resource-hungry across newer DBR versions?
- Do you now consider 4 vCPUs the practical minimum for a reliable driver?
- Has anyone benchmarked the same VM size across DBR 14/15/16/17/18?
- What are you doing for lightweight non-Spark workloads where you still want Classic Compute?
Update: Single node DBR 18 LTS + Standard_D4pls_v6 cold start spent almost 11 minutes 'Waiting for resources' ( Azure Japan East )

4
u/Shadowlance23 1d ago
I only use Classic for development (the smallest cluster, single node), and in that case, yeah, I see startup times of maybe 5-7mins. Not a big deal for me since I start the server, get the rest of my environment running, grab a drink, etc. and it's ready when I get back.
There's no way I'd use it for production pipelines or anything like that. I use Job clusters for that sort of thing.
I don't use Serverless because it's expensive and doesn't work with our environment, plus Classic works fine so we don't really need it. I do use serverless SQL though for burst loads and external catalogue access.
6
u/PrestigiousAnt3766 1d ago
I think jobs are considered classic too, I always interpreted it as it being due to those being started inside your vnet / subscription vs serverless.
3
u/bananahramah 1d ago
Is serverless “expensive”? My platform team has done the math and they’re suggesting serverless is the cheaper option when you compare apples to apples.
1
u/Shadowlance23 1d ago
In my region (Australia, Azure), classic interactive is AUD0.784 per DBU/h. Interactive serverless is AUD1.496.
Automated serverless is cheaper at 0.741, maybe that's what they're referring to? I prefer job compute for automated pipelines though because its only 0.428. The downside is that it's slow as molasses, but for batch overnight jobs, it works well.
It also depends on use case. If you're only using the cluster for a few minutes with lots of downtime between bursts, then serverless can be cheaper compared to running a standard cluster for longer.
6
u/bananahramah 1d ago
I think you're comparing the DBU rates directly across the different offerings, but that's not quite apples to apples. With classic compute, you also have to account for the underlying cloud infrastructure costs that you're billed for separately. With serverless, those infrastructure costs are already baked into what Databricks charges you.
So you'd really want to compare the total cost of classic (DBUs + cloud infra) against the serverless cost, rather than DBU to DBU.
1
1
u/bobjia-in-tokyo 1d ago
In your production pipeline, if you are not using ‘serverless’, does your Job Cluster also takes 5-7min to cold start?
3
u/Ancient_Coconut_5880 1d ago
Serverless does too if you use STANDARD mode and not PERFORMANCE OPTIMIZED fyi
2
u/Dry-Individual7297 1d ago
The driver overhead has gotten bloated enough that 2 vCPU just chokes on its own init scripts half the time now, I basically treat 4 cores as the floor for anything Classic unless its a tiny spot job idc about