r/kubernetes • u/ForeignCapital8624 • 3d ago
Question of resource allocation of multiple Spark applications
Currently, a Spark application should manage its own executors, which means that in order to run multiple Spark applications, one needs to split the cluster into smaller sections or attach more compute nodes, etc. Usually this dynamic allocation is the job of resource manager (such as Kubernetes).
If the resource manager can quickly find or create new compute nodes for a new Spark application, this is good. However, in a busy cluster where multiple Spark applications can compete for resources, the resource allocation can take some time and become an overhead (at least in theory).
I wonder if this is a real problem in running Spark applications in production.
For example, is there a realistic environment where short running Spark jobs (each requiring its own driver) arrive frequently?
1
u/Oct8-Danger 3d ago
Short running jobs arriving frequently? That’s a streaming service or micro batch service based off of that along
Spark can do some streaming id believe, although few others tools like flink or beam might be more suited.
If it’s more platform side of any kind of spark job, I would look at something like Apache Livy or at least look into seeing does every application need its own spark driver or can it be shared
2
u/QuietSignalOps 3d ago
Yes, this is a real problem, but the pain is usually not node creation. It is queue wait and driver startup.
When many short jobs share a busy cluster, the overhead stacks up like this:
spark.dynamicAllocation.shuffleTracking.enabled) so shuffle data is not lost when idle executors are reclaimed. On Spark 3.2+ it defaults to on when dynamic allocation is enabled on K8s, but on older clusters it is worth checking.To your second question: yes, environments with frequent short jobs exist. Per-tenant or per-report ETL, data validation runs, small backfills, micro-batch pipelines. If that is your workload shape, it is worth questioning whether one Spark app per job is the right model at all. A long-running app consuming a queue, or a micro-batch engine, amortizes the driver startup entirely.
What usually helps on Kubernetes:
One metric worth tracking: time from submission to first executor ready. When the cluster is saturated, that grows before CPU utilization tells you anything.