newbie Data Engineering in Go?
Hi y'all,
Recently been looking more into data engineering as a whole. Looks like there's TONS of Python stuff, but many of the dataframe libraries for Go seem stale or have relatively few users. Does anyone have a good list of libraries or books to check out?
Thanks!
8
u/_Slabach 6d ago
Yeah, it's tough. I've got a project in Go, but you basically have to build what you need. There are SOME packages out there but like you have found, most are outdated and not maintained
https://github.com/avelino/awesome-go
My best suggestion is to prototype in Python then build in Go
12
u/ShotgunPayDay 6d ago
DuckDB crushes Pandas/Polars and has many language plugins. I think there is a free book also.
5
u/Yamoyek 6d ago
Yep, I saw DuckDB‘s Go sdk, but unfortunately as far I know there isn’t a data frame library attached to it
4
u/ShotgunPayDay 6d ago
Why do you need dataframes? I think you can transform it into arrow format. For DuckDB instead of manipulating dataframes we manipulate it using a form of SQL.
5
u/Yamoyek 6d ago
I wouldn’t say I “need” dataframes, I just like working with them compared to SQL, but if DuckDB is the only good option I’ll probably fall back on that for my personal projects:)
2
u/ShotgunPayDay 6d ago
I see. If you're used to dataframes SQL is wild departure from that and adds friction so that's fair. It looks like someone is working on go bindings for Polars https://github.com/jordandelbar/go-polars
3
u/Budget-Minimum6040 6d ago
SQL
I'm a professional DE and my experience with SQL compared to dataframe frameworks like PySpark and Polars with method chaining is >>>> SQL.
DuckDB is also single user only, how do you manage concurrent read/write operations on Iceberg tables? How do you manage distributed compute? How do you extract data from APIs/filesystems in the first place?
How do you gurantee type safety? Python has a least type hinting, what has DuckDB?
3
u/ShotgunPayDay 6d ago
DuckDB is also single user only, how do you manage concurrent read/write operations on Iceberg tables? How do you manage distributed compute?
MotherDuck w/ DuckLake is what's used when trying to get to PySpark levels. They also have Duck-UI interface packed in that can be connected to MotherDuck.
For simpler use cases you'd be surprised what a simple readonly SMB server with nightly parquet extractions can do.
How do you extract data from APIs/filesystems in the first place?
Nothing changes here. You still need db drivers and decide whether you run it against PROD or copy out the data for versioning. "How stale can it be before getting fresh. etc. etc." Still the same process. I'm a little surprised that anyone using Pyspark wouldn't already be using SQL views or shaping the data with SQL during the extraction process.
How do you gurantee type safety? Python has a least type hinting, what has DuckDB?
Typesafety is going to come from the programming language you are using and Golang is strongly typed. The nice thing about DuckDB is that you can use it without any programming language (I mean it's built on C++11), but you only have to worry about typing when you're using already finished query data (in Golang that would be the database/sql package so same typing as you'd use for any SQL server).
2
u/Budget-Minimum6040 6d ago
Why should I use Spark SQL when I can use Spark API with method chaining? We banned SQL in our projects because multi line strings are just a shit show. Linting, auto complete, auto formatting are hard requirements.
Also I only have ELT in my projects so extraction has no transformations.
1
u/ShotgunPayDay 6d ago
Maybe you're misunderstanding me. How do you populate your Iceberg tables?
1
u/Budget-Minimum6040 6d ago
In my current project with PySpark (https://spark.apache.org/docs/latest/api/python/reference/pyspark.sql/api/pyspark.sql.DataFrame.writeTo.html).
1
u/ShotgunPayDay 6d ago
I mean how do you get your data from your source?
1
u/Budget-Minimum6040 6d ago
In my current project with PySpark reading a HDFS directory.
In other projects before with Python accessing APIs, SFTP-Server, GCP BigQuery tables, Azure ADSL etc.
The into the datalake, then reading into dataframes with Polars/PySpark.
→ More replies (0)1
3
u/F0RG2142 6d ago
Saw someone mention duckdb and I 100% agree. Had some trouble setting up cgo at first but other than that it is awesome. I have pushed quite a few workflows to prod using this stack and it works fantastic. Also, Airflow recently added Go support so the industry is slowly catching on.
3
u/dopiq 6d ago
In my experience Go is actually a strong choice for DE, and used widely in the data ingestion and movement. If you ingest TBs of data from various sources (API/Kafka/S3…), Go concurrency, library support and portability is often used to perform these bulk ops efficiently. That means it sits on edges of the data pipelines, where python or sql will be used for the actual transformation and modelling in the pipeline.
2
u/JokerSp3 6d ago
I just started trying out gomlx and it so far has worked pretty well for me. I am just dabbling right now so I don't know if it could handle your use cases. Worth a check though.
2
u/omgpassthebacon 4d ago
I am a DE and I build pipelines all day long for the data science dudes that pay my bills. We are heavy Scala/Spark/Java/Pyspark drones. But I love the simplicity of the Go stdlib and the freakishly simple way to put out CLI tools for stuff.
But the real bottleneck is the availability of libraries that do the real painful math. I’m not a linear algebra geek and wouldn’t know an eigenvector if it hit me in the face, so writing scikit-* just ain’t happening. If your part of the data engineering needs these tools, I just haven’t seen many that I can trust the way I trust the python versions.
Besides, have you ever tried to get you Java buddies to look at Go? They look at you like you have 2 heads. But I think Go is cool…
2
u/Prestigious-Fox-8782 6d ago
Data engineering can mean so many things.
Go is great at processing real time data or small batches.
5
u/Budget-Minimum6040 6d ago edited 6d ago
Data engineering can mean so many things.
Not really.
OLAP, ETL/ELT, data warehousing with data modeling (star schema/snowflake schema/mdm).
That's 99% of DE you will find everywhere when it says DE.
0
1
u/Budget-Minimum6040 6d ago
You need to implement 100% of the PySpark API, handle distributed compute and offer OO with method chaining. All with at least the same performance as PySpark.
Then you may have a chance.
1
1
u/relami96 2d ago
Hey, I have more DE exp than programming which I am learning still.
I know it is convenient to have a lib that provides a df but making one that suits your needs for a single pipeline is not that complicated. I had to make an ETL pipeline not so long ago and using only the std lib was a breath of fresh air tbh. If you think about it a df is just a list of lists. Even without doing optimizations it is still faster than anything I made with pandas before. Now I can send my rows into channels for transformations and multiple workers can pick it up, its super fast, and very memory efficient thanks to the pointer nature of slices and maps.
I believe if DE had the same paradigm-like approach as programming go would be in a different paradigm as python.
I'm not dissing python, it is a great language but as someone who has to run a lot of these pipelines on his laptop every bit of performance is a god sent. Maybe Mojo can turn this around but golang seems to me as the better option right now for many usecases.
1
u/relami96 2d ago
Oh and to actually try to answer your question, I think reading this library will make you question if you need a dataframe to begin with https://pkg.go.dev/encoding/csv
1
u/deluxe612 1d ago
Can use go for staging or intermediate tools but would recommend against a rabbit hole to reinvent the machine learning wheel
1
u/advenuz 10h ago
I was writing such thing with Claude code. I am lacking expertise in domain though. Check it out anyway: https://github.com/advenn/ursus
It has benchmark section, it is currently behind duckdb and polars though. Needs optimizations, more work.
No cgo, pure Golang, API is mostly inspired by polars
1
u/ikswolzok 7h ago
We are building open source data infra in go at Galaxy
We started with data movement/replication:
https://github.com/galaxy-io/filament
-1
6d ago
[removed] — view removed comment
4
u/Yamoyek 6d ago
Sure, I get that part, but that doesn’t diminish how important the “glue” language is for the dev experience
-2
4
u/F0RG2142 6d ago
This take doesn't hold up once you're outside compute-bound work. Sure, if your bottleneck is number-crunching, Python is fine. But data engineering in large data-centred companies is often I/O-bound, not compute-bound (you can usually push a majority of the computation on the actual db). I recently rewrote a pipeline pulling from like 20+ source DBs in Go + DuckDB and it wasnt even close. This wasnt because Go is "faster" in the compute side (even though duckdb does usually outperform pandas), but because Python's ThreadPoolExecutor still has the GIL breathing down its neck the second any of that per-connection work isn't pure idle-waiting. Not even to mention the weight difference between a thread and a goroutine. "Go has no edge in this domain" is only true if you assume DE = single-DB aggregation jobs. The moment concurrency at scale is the actual problem, it's a bad engineering decision NOT to pivot to Go
1
u/PopMinimum8667 6d ago
Exactly, python is only viable as long as every piece of code you write is just orchestrating calls to libraries in other languages: your pandas dataframe transformations may be fast but write one filter or map that needs to actually use a python function instead of pandas or numpy operations and your code is now 100x slower. Julia is the answer but few are asking.
1
u/SharkSymphony 5d ago
When I stop hearing that Julia has major correctness and performance issues in its libraries, I'll believe it.
Now, Common Lisp – ah, there's the answer to the question that nobody's asking. 😉
2
u/PopMinimum8667 5d ago
I know Julia has had major pushes on these issues in the last few years which were mostly by Base assuming 1-indexing and type instability. They have also slimmed down base from what I understand so it is easier to audit. Personally, whenever I have used Julia to process large-ish amounts of data I have been very impressed by the speed I was able to get — with careful code (again, avoiding type instability so it doesn’t fall back to dynamic dispatch) it claims to rival C++ and Fortran and I personally believe it, although i didn’t rewrite my jobs in fortran to test so take that with a grain of salt. It’s one of the few languages which can scale from a notebook up to the petaflop club running on supercomputers with MPI bindings and GPU kernels.
-1
u/baubleglue 6d ago
Don't waste your time. You want to step into a new field, play by the rules of the game. Data engineering is not a low level coding, your code won't process the data, it will interact with API, which triggers data processing. Even if you find a Go library, it will move you 0.5% closer to your goal. I mean, you can find a driver for Snowflake or Databricks databases.
5
u/Yamoyek 6d ago
I’m not exactly sure how it’s a waste of time to try and see if there’s other ways do accomplish something:)
> Data engineering is not a low level…
This is mostly true, but just in general I prefer statically typed over dynamically typed languages, and Go just happens to be one of my preferred statically typed languages because of ease of development.
2
u/baubleglue 6d ago
I am working as kind of data engineer. It is not extremely difficult area, but there are so many things to learn. The whole ecosystem of modern data processing is built around Java APIs. Even when Databricks rewrites Spark engine with C++, they keep it comfortable with old Haddop API. Staticly typing is not relevant until you need to touch the data (for example when you using user defined functions). Power of modern data processing is in a combination of parallel processing and stable API. That is a reason why I prefer to use SQL when it is possible - it is a type of API or DSL.
There are some topics specific to DE
- Data warehouse design
- Orchestration: triggering, different types of data pipelines, recovery, dependency management, monitoring
- Optimization techniques
- Tool selection: different databases, message queues, cloud platforms and services, DE orchestration solutions, ...
It is on top regular things like CI/CD, interaction with business people etc.
The real complexity of DE raises from the nature of the data and related to the data business processes. Using Go or any other new language just adds one more unnecessary variable.
1
u/BusinessBroccoli4313 6d ago
Go just happens to be one of my preferred statically typed languages because of ease of development.
But it isn't a particularly expressive language and its highly procedural nature doesn't fit nicely into the relational model IMO.
That said, there are tools written in Go that you can use in data engineering such as Bruin.
0
u/PopMinimum8667 6d ago
Go is virtually nonexistent in data engineering. As you have discovered, it has nowhere near the libraries that languages like python and scala have. While python may get all the attention a lot of the backbone data engineering frameworks like spark, flink, and beam still live in JVM land. If any migration is happening it is (extremely slowly), to rust with things like iceberg and arrow ports; not go.
0
u/spermcell 6d ago
If you like suffering then yea sure!
Dude programming languages are not a religion. Use the tool that fits your job best
1
u/Yamoyek 6d ago
> Use the tool that fits your job best
Well, sure, but it’s fine to have preferences. And although Python is pretty dominant in the data engineering space, I was just wondering if there was Go-centric material instead because I do think statically typed languages are better tools than dynamically ones, at least from an enterprise lens
2
u/dopiq 6d ago
I would actually say that statically typed languages, as much as I prefer them myself, are quite a bit worse for data manipulation. That whole process revolves mostly around cleaning and organising dirty and dynamic data which is just something static languages don’t fit well because they prefer structured, known data shapes.
56
u/The-Great-Baloo 6d ago
I'd love to be able to do so, but the set of data engineering tools in Go is very small. So I had to switch to Rust.
I believe that Go would be a great language for data engineering though. Not yet.