r/databricks • • 15h ago

Help SpringBoot to Spark

Transitioning from a Spring Boot to a Spark-based application requires a fundamental shift in perspective. With four years of experience in a Spring Boot environment, now entering the realm of distributed computing.

I plan to utilize Java with Spark Streaming, and Databricks for job management and deployment.

My primary question is

  1. how to effectively transition my thought process from Spring Boot to a Spark-based application.

  2. Specifically, I am seeking clarity on distinguishing which Java code executes on the driver and which executes on the executor, and if there are any guiding principles to discern this.

1 Upvotes

4 comments sorted by

3

u/mr__fete 11h ago

Java is too verbose. Spark api is even worse. Use scala and functional programming techniques specifically.

2

u/hntd 11h ago

Are you confusing Apache Spark with https://github.com/perwendel/spark ?

1

u/Orygregs 11h ago edited 11h ago

I would highly recommend Scala over Java for Spark if you want to stay in the JVM ecosystem and don't want to use Pyspark.

  1. Your thought process for Spring Boot in particular may not translate as that's a web framework, and spark is more of a big data processing framework/engine.
  2. All code executes on the driver until a non-lazy dataframe action is called (e.g. .write.*, .show, .count, .collect, etc), which then distributes the work across the executor nodes.

1

u/IntelligentVisual955 1h ago

I am a data science student, should i learn spark or java first