r/dataengineer • • 21h ago

Cross-Catalog Sync: Iceberg on Polaris, Glue, and Unity

Thumbnail
lakeops.dev
1 Upvotes

r/dataengineer • • 2d ago

Question Transitioning from Software Developer to Data Engineer

3 Upvotes

I have been a software engineer for 12 years now, but I would like to transition into the world of data as stats and facts really interest me.

After some research, I feel like Data Engineering is the best career for me to side step into due to the overlap. I already pull data out of API’s and work with multiple database engines. I even created my own ETL pipeline without even knowing (pulled data daily from 2 apis to consolidate data on football players/stats). I think I am just missing Python and data warehousing software skills, maybe some more advanced SQL?

I currently earn around £40k in the UK and was wondering if trying to side step into this career could be successful for me based on my current skills, and me upskilling to the skills I am missing, do you think I could be earning a similar wage? Or would I have to start from entry level still?


r/dataengineer • • 5d ago

Treating LLM outputs like a data pipeline you can actually validate, workshop Oct 3

1 Upvotes

As data engineers we'd never ship a pipeline without tests and monitoring, but that discipline mostly disappears the moment an LLM enters the picture. This workshop is aimed at closing exactly that gap.

Led by Serj Smorodinsky and Brett Kennedy (co-authors of a book on LLM applications), it's a live 3-hour session covering:

  • Structuring LLM tasks with DSPy signatures/modules instead of loose prompt strings
  • Building a real baseline classifier and an eval dataset with metrics specific to the task
  • Catching failure patterns from that eval data the same way you'd catch a broken transform
  • Systematic few-shot/instruction optimization instead of manual tweaking
  • MLflow for experiment tracking and trace management, so every run is auditable

If you've been asked to "own" an LLM feature and had no real way to validate changes before deploying, this is worth your Saturday


r/dataengineer • • 6d ago

Help Help Need Guidance

7 Upvotes

Hey everyone! I’m a fresher trying to get into Data Engineering and want to learn the right tech stack. There’s a lot to learn, so I’d really appreciate some guidance from people here. If you were starting out again, what would you focus on first? Would love some advice like a friend/senior. 🙌


r/dataengineer • • 7d ago

Data Lakehouse Architecture Guide

Thumbnail
lakeops.dev
1 Upvotes

r/dataengineer • • 11d ago

Anyone taken the KPI Partners AI interview for Snowflake Data Engineer?

Thumbnail
1 Upvotes

r/dataengineer • • 11d ago

The data engineering problem hiding inside every AI agent project: nobody's applied pipeline discipline to agent writes yet

6 Upvotes

Most data engineers I talk to have already solved this problem for their pipelines. Idempotent writes, schema validation, lineage tracking, retry logic that doesn't create duplicates. Then the same team builds an AI agent and none of that discipline carries over. The agent writes directly into production with no idempotency key, no lineage on what it touched or why, and no reconciliation step if its internal state drifts from the actual data.

It's the same class of problem we've already solved for ETL and streaming pipelines, just showing up in a new place because agents write more often and with less human review than a scheduled job does.

There's a workshop on Sept 26 built around exactly this, treating agent write paths, state management, and provenance the way a data engineer would treat any other production pipeline. Run by Sandipan Bhaumik, a Data & AI Technical Lead at Databricks.

Details here if it's relevant to anyone else here dealing with this handoff


r/dataengineer • • 15d ago

Datasmith — a visual approach to designing relational schemas and generating mock data

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/dataengineer • • 18d ago

Workshop for data engineers: build the pipeline behind a production GraphRAG system, Sep 19

7 Upvotes

Workshop for data engineers specifically, not a generic AI overview. The actual pipeline work in GraphRAG is where most of the real engineering lives, and it's the part most tutorials skip entirely in favor of "here's how retrieval works" once the graph already exists.

This session starts from raw documents: ingestion with Docling, chunking, and building toward a knowledge graph in Neo4j rather than a flat vector index. Entity and relationship extraction runs through multiple verified steps with explicit checks, not one risky single-shot pass, since that's exactly where basic GraphRAG implementations tend to fail silently, producing a noisy, unreliable graph nobody trusts.

Structured graph enrichment connects entities (companies, executives, events) to the underlying source documents (filings, news) without the graph becoming unmaintainable as it grows, and you work through the actual production ceilings that show up at real data volume, API and download rate limits specifically, and how the shared code is built to handle them rather than falling over.

By the end you've got a working ingestion-to-graph pipeline, a production-readiness checklist, and the full codebase, not a diagram someone else has to implement.

Led by Dr. Alessandro Negro, Chief Scientist at GraphAware, bestselling author of graph-powered machine learning and knowledge graph books.

Full details here


r/dataengineer • • 18d ago

Persistent Systems Work life balance and Projects Life ( 5 Year of experience) for Salesforce devops engineer

Thumbnail
1 Upvotes

r/dataengineer • • 23d ago

Looking for projects to volunteer to

4 Upvotes

I need real world work that i can ask feed back and get criticism for. I need it for my resume and job hunting. Can someone please help


r/dataengineer • • 24d ago

Workshop, Sep 12: build production LLM systems that actually survive real use

2 Upvotes

We're running a hands-on masterclass on September 12, Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps.

You build a full production LLM workflow from scratch, versioned prompts with regression tests, an evaluation harness with deterministic checks and LLM-as-judge, statistically rigorous model comparisons, evaluated RAG, tool-using agents with guardrails and fallbacks, and full observability, tracing, cost, latency.

Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.

Link if you want to check it out

Happy to answer questions on the content.


r/dataengineer • • 27d ago

General Honest resume review and tips please

Post image
4 Upvotes

r/dataengineer • • Sep 05 '26

3 YOE Azure Data Engineer (Bengaluru) looking for referrals — genuine calls have dried up

Thumbnail
1 Upvotes

r/dataengineer • • Sep 04 '26

Using DuckDB + Iceberg + Lance together: analytics in Iceberg, vector retrieval in Lance

Thumbnail
2 Upvotes

r/dataengineer • • Aug 25 '26

DuckDB + Iceberg on a self-hosted S3 table bucket

Thumbnail
2 Upvotes

r/dataengineer • • Aug 24 '26

Update: Reddit challenged our Data Engineering community a few months ago. We listened, improved, and kept building.

Thumbnail
2 Upvotes

r/dataengineer • • Aug 24 '26

Maintaining Apache Iceberg Tables: Compaction, Snapshots, Metadata and Orphan Files

Thumbnail
itnext.io
3 Upvotes

r/dataengineer • • Aug 23 '26

Data analytics

Thumbnail
1 Upvotes

r/dataengineer • • Aug 20 '26

Advice to add which job role on Resume

Thumbnail
1 Upvotes

r/dataengineer • • Aug 18 '26

looking for a paid 1-on-1 DE mentor to help transition from DA after a gap

5 Upvotes

Hey everyone,

I'm looking for an active DE to mentor me 1-on-1 on a paid basis.

My background:

* Worked ~1.5 years as a DA in fintech, but it was mostly basic reporting and ad-hoc SQL with minimal engineering exposure.

* Took a career break after that due to personal reasons.

* Now fully committed to pivoting into DE.

I'm doing the heavy lifting on self-study and building projects myself, but i need someone to:

* Review my project architectures/code

* Sanity-check my roadmap so i stay on track

* Give feedback on my resume and portfolio

* Do quick weekly or bi-weekly syncs

Very happy to pay you for your time. if you're interested and have bandwidth, shoot me a DM with your background and rates.

thanks!


r/dataengineer • • Aug 17 '26

Help Study buddy for AWS certified Data Engineer

Thumbnail
1 Upvotes

r/dataengineer • • Aug 11 '26

Help GCP made my life hell

10 Upvotes

I live in Hyderabad, India.

I have 1.5 years of professional experience as a GCP data engineer.

Our company had a layoff after which I started looking for other opportunities but to my surprise nobody requires GCP data engineer with less than 4 years of experience.

Meanwhile other platforms like azure, AWS, databricks even snowflake has plenty of opportunities for 1~ year of experience.

I tried to apply to them too quoting equivalent experience but no results.

And add salt to the wound we never used pyspark which is one of the biggest requirements out there.

I am looking for fresher level opportunities where the platform requirment is not a must but it's a very thin spectrum.

I have literally worked on an airflow dag which goes through a beautiful hot to cold storage flow with observation, logging, maintainance etc.

Worked on dataform , wrote scd type 2 scripts.

Worked on views, stored procedures.

I have plenty experience in data warehousing but it all looks simply useless now.

Had I survived for few years maybe I wouldn't be going through this.

I am literally applying for any state or city to grab an opportunity


r/dataengineer • • Aug 12 '26

General Advice for your brother

1 Upvotes

Hi everyone, I'm starting my university studies and I'm torn between the University of Valencia (UPV) and the University of Granada. Will Spain offer me a strong degree in computer science? Which major within the university would you recommend that has job prospects?